SailPoint

Senior Staff DevOps Engineer

SailPoint · Canada · $142K–$240K/yr

Software Development · 1,001-5,000 employees

19 h ago
Remote Principal (10+ yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead the architecture, implementation, and operation of a large-scale service mesh across hundreds of microservices on AWS EKS. Define and enforce platform standards, mentor engineering teams, and ensure infrastructure compliance with PCI DSS requirements.

What they look for

Kubernetes AWS Service mesh Istio Linkerd Terraform Python Go CI/CD GitOps Microservices PCI DSS Linux Distributed systems EKS Infrastructure as code

Requirements

Requires over 10 years of experience in production operations and containerization, with at least 5 years of hands-on experience in Kubernetes and Terraform. Candidates must demonstrate expertise in scaling service meshes and operating highly available SaaS environments.

Benefits

Medical insurance Dental insurance Vision insurance Short-term disability Long-term disability Life insurance Accidental death and dismemberment insurance 401(k) savings and investment plan Flexible vacation policy Paid holidays Sick leave Paid parental leave Employee assistance program Legal assistance Critical illness insurance Accident insurance Hospital indemnity Pet insurance Health savings account

Full description

About SailPoint 

SailPoint is the leader in identity security for the cloud enterprise. Our identity security solutions secure and enable thousands of companies worldwide, giving customers unmatched visibility into their digital workforce and ensuring workers have the right access, no more and no less. 

Built on a foundation of AI and machine learning, our Identity Security Cloud (Atlas) platform delivers the right level of access to the right identities and resources at the right time, matching the scale, velocity, and changing needs of today's cloud-oriented enterprise. SailPoint is all-in on the AI revolution, and our engineers have access to the latest frontier models and agentic frameworks. 

 

About the Role 

As a Senior Staff DevOps Engineer on the Infrastructure Platform team, you will be a technical leader responsible for designing, building, and operating SailPoint's global Identity Security Cloud infrastructure on AWS. You will partner with engineering teams across the US, India, EMEA, and APAC to deliver a resilient, secure platform. 

You will serve as the Kubernetes platform leader for the team, driving large-scale improvements such as implementing a service mesh (Istio, Linkerd, or AWS App Mesh) across hundreds of microservices and production EKS clusters, and guiding engineers on cloud-native deployment patterns. You will also play a key role in supporting SailPoint's PCI compliance initiative, helping ensure our platform infrastructure meets PCI DSS requirements. 

This is a fully remote position for candidates based in the USA or Canada. The role carries significant technical influence across the organization, setting architecture direction, mentoring engineers, and driving operational excellence without requiring people management. 

The ideal candidate is a self-starter who thrives in complex, fast-paced SaaS environments, has proven hands-on experience rolling out and operating a service mesh at scale across large microservices environments, brings expert-level Kubernetes and AWS cloud knowledge, values quality and reliability, and is excited to work on an infrastructure team pushing the boundaries of cloud-native operations. 

About the Team 

The Infrastructure Platform Team designs and operates the foundational cloud infrastructure that powers SailPoint's Identity Security Cloud. Our platform runs on a mature, cloud-native, event-driven microservices architecture on AWS, with Kubernetes, GitOps, and Infrastructure-as-Code at the core. We also build and operate the data platform and AI/ML tooling that underpin SailPoint's identity intelligence capabilities, supporting large-scale data pipelines, model serving infrastructure, and the integrations that bring AI-driven insights into our products. 

The Infrastructure Platform Team designs and operates the foundational cloud infrastructure that powers SailPoint’s Identity Security Cloud. Our platform is built on a combination of microservices, data pipelines, machine learning, and AI tooling, most of which relies on Kubernetes for compute. 

Engineers work closely with global peers with a follow-the-sun on-call rotation. The team owns all of SailPoint's cloud infrastructure, compute capabilities, and operational practices for a mission critical, 24x7 SaaS environment serving large enterprises worldwide. We focus on security, modern practices, and managing web platforms that serve millions of concurrent requests. 

 

Key Responsibilities 

  • Own and lead the full lifecycle of service mesh implementation at scale, from architecture and rollout planning through production operations across hundreds of microservices running on EKS. 
  • Establish service mesh standards and governance adopted organization-wide, including onboarding runbooks, sidecar injection policies, traffic policy templates, and documented failure-mode playbooks for production incidents. 
  • Mentor and upskill engineers across teams on service mesh architecture, troubleshooting at scale, performance tuning, and capacity planning for mesh-heavy environments. 
  • Design, operate, and optimize production Kubernetes clusters at scale on AWS (EKS), including cluster architecture, upgrades, node management, networking, storage, and multi-tenant isolation patterns. 
  • Define and drive Kubernetes standards and best practices across teams, including workload design, resource management, security hardening (RBAC, Pod Security Standards, network policies), Helm/chart conventions, and deployment patterns. 
  • Design and scale infrastructure to meet rapidly increasing customer demand, data sovereignty requirements, and regional expansion. 
  • Automate deployment, monitoring, incident response, and capacity management using GitOps and CI/CD best practices. 
  • Develop and improve operational practices, runbooks, and platform engineering standards. 
  • Collaborate with development teams to bring new features and services into production safely and efficiently. 
  • Proactively meet information security and compliance standards (e.g., PCI DSS), including supporting SailPoint's PCI compliance initiative through secure platform design, controls implementation, and audit readiness. 
  • Participate in and help improve the on-call rotation; drive post-incident reviews and systemic fixes. 

 

Background and Experience 

Required 

  • Strong interpersonal and teaming skills, with the ability to set and enforce process and influence engineers across teams and geographies. 
  • Ability to operate effectively in an agile, entrepreneurial environment with global stakeholders. 
  • Prior experience as a technical lead or Staff+ IC in a global engineering organization. 
  • 3+ years of hands-on experience designing, implementing, and operating a service mesh at scale in production Kubernetes environments. 
  • 10+ years of experience in 24x7 production operations, supporting highly available SaaS or cloud service environments. 
  • 10+ years of experience with containerization, virtualization, and configuration management technologies. 
  • 5+ years of hands-on experience with Kubernetes in production at scale. 
  • 5+ years of experience with Terraform (IaC), managing infrastructure across multiple AWS accounts and regions. 
  • 5+ years of experience designing and implementing CI/CD pipelines, especially for Terraform, Kubernetes, and microservices. 
  • 5+ years of experience with scripting/programming languages (Python, Go, or similar) and strong shell scripting proficiency. 
  • Strong understanding of Linux, networking, distributed systems, and production troubleshooting. 
  • Demonstrated experience scaling a service mesh across a large microservices fleet, including phased adoption strategy, sidecar resource management, control plane scaling, and performance tuning under high request volume. 
  • Experience with monitoring and logging stacks (e.g., Prometheus, Grafana, OpenSearch or equivalent). 

Preferred 

  • Experience with multi-cluster or multi-region service mesh federation (e.g., Istio multi-cluster, Linkerd multicluster extension) in production. 
  • Familiarity with compliance frameworks in regulated enterprise SaaS, including PCI DSS and FedRAMP-adjacent practices, with experience implementing or operating platforms subject to PCI compliance requirements preferred. 

 

What Success Looks Like 

First 30 Days 

  • Onboard into the role; learn Identity Security Cloud architecture, platform tooling, and team processes. 
  • Build relationships with peers and stakeholders across the US and global DevOps/SRE teams. 
  •  
  • Begin scoping and designing the enterprise-scale service mesh rollout strategy.Join team ceremonies, contribute to in-flight projects, and begin participating in on-call shadowing. 

First 90 Days 

  • Own significant platform initiatives and contribute to architecture decisions and operational improvements. 
  • Deliver a prioritized service mesh scale and maturity roadmap, covering observability gaps, security policy coverage, and onboarding backlog.
  • Complete at least one major service mesh scaling milestone, such as full mTLS enforcement across a production cluster tier, mesh-based progressive delivery adoption, or control plane high-availability hardening.
  • Join the on-call rotation; become proficient in core services, escalation paths, and runbooks.
  • Mentor engineers and document at least one new platform pattern or operational standard.

First 6 Months

  • Lead end-to-end service mesh adoption and standardization at scale, including full mTLS coverage, traffic policy governance, mesh-native observability dashboards, documented failure playbooks, and onboarding patterns adopted organization-wide.
  • Become a recognized SME on platform services and independently handle complex production escalations.
  • Influence roadmap and engineering practices across the broader Infrastructure organization.

First 12 Months

  • Establish durable standards, guardrails, and documentation adopted by multiple teams.
  • Demonstrate sustained impact on uptime, deployment success rate, cost efficiency, or engineering productivity.

Work Model

This is a fully remote position for candidates based in the USA or Canada. Candidates must be eligible to work in the United States or Canada and be available for collaboration with global teams across US, EMEA, and APAC time zones. Participation in an on-call rotation is required.

We anticipate approximately 10% travel, with flexibility to travel to our headquarters in Austin, TX as needed for leadership meetings, architectural planning sessions, and cross-functional team discussions. Candidates should be comfortable representing their team's technical direction in person during these engagements.

Education

Bachelor's and/or Master's degree in Computer Science or equivalent technical experience.

Benefits and Compensation listed vary based on the location of your employment and the nature of your employment with SailPoint.

As a part of the total compensation package, this role may be eligible for the SailPoint Corporate Bonus Plan or a role-specific commission, along with potential eligibility for equity participation. SailPoint maintains broad salary ranges for its roles to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect SailPoint’s differing products, industries, and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. We estimate the base salary, for US-based employees, will be in this range from (min-max, USD):

$142,500 - $240,248.00Base salaries for employees based in other locations are competitive for the employee’s home location.

Benefits Overview

1. Health and wellness coverage: Medical, dental, and vision insurance

2. Disability coverage: Short-term and long-term disability

3. Life protection: Life insurance and Accidental Death & Dismemberment (AD&D)

4. Additional life coverage options: Supplemental life insurance for employees, spouses, and children

5. Flexible spending accounts for health care, and dependent care; limited purpose flexible spending account

6. Financial security: 401(k) Savings and Investment Plan with company matching

7. Time off benefits: Flexible vacation policy

8. Holidays: 8 paid holidays annually

9. Sick leave

10. Parental support: Paid parental leave

11. Employee Assistance Program (EAP) and Care Counselors

12. Voluntary benefits: Legal Assistance, Critical Illness, Accident, Hospital Indemnity and Pet Insurance options

13. Health Savings Account (HSA) with employer contribution

SailPoint is an equal opportunity employer and we welcome all qualified candidates to apply to join our team.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other category protected by applicable law.  

Alternative methods of applying for employment are available to individuals unable to submit an application through this site because of a disability. Contact applicationassistance@sailpoint.com or mail to 11120 Four Points Dr, Suite 100, Austin, TX 78726, to discuss reasonable accommodations.  NOTE: Any unsolicited resumes sent by candidates or agencies to this email will not be considered for current openings at SailPoint.