Principal Site Reliability Engineer
UKG Seattle, Washington, United States · $184K–$265K/yr
Software Development · 10,001+ employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
Define and evolve the organization's Site Reliability Engineering strategy, vision, and technical roadmap in partnership with leadership. Lead cross-functional initiatives to establish reliability standards, mentor SRE talent, and drive system performance and availability at scale.
What they look for
Requirements
Requires 12+ years of experience in software or systems engineering with at least 10 years in public cloud environments. Candidates must demonstrate expertise in distributed systems, observability, and modern CI/CD practices alongside strong leadership and communication skills.
Benefits
Full description
Why UKG:
At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our workforce operating platform. Helping people get paid, grow in their careers, and shape the future of their industries. That’s what we do.
We never stop learning. We never stop challenging the norm. We push for better, and we celebrate the wins along the way. Here, you’ll get flexibility that’s real, benefits you can count on, and a team that succeeds together. Because at UKG, your work matters—and so do you.
About the Team Principal Site Reliability Engineers (SREs) at UKG are strategic technical leaders who play a critical role in shaping the reliability, scalability, and performance of our organization's services, at scale. They bring deep expertise across service delivery, infrastructure, and cloud architecture, and apply software engineering principles to solve complex operational challenges. Principal SREs drive technical vision and strategy across multiple teams and business units. In this role, you will define the operational reliability strategy, architect large-scale solutions, and establish the standards and practices that enable teams across UKG to build and operate highly reliable services. This includes guiding the evolution of CI/CD ecosystems, designing automated testing frameworks, building capacity planning and performance analysis systems, establishing observability standards, and creating reliability automation strategies as self-healing and remediation at scale. Principal SREs are passionate about driving business outcomes through technical excellence. They shape operational culture around reliability, mentor and develop SRE talent across teams, and relentlessly pursue the highest standards of customer experience through strategic automation and process innovation. This is a principal-level individual contributor and technical leadership role, focused on operational reliability strategy, cross-functional influence, and enterprise-scale impact.
Principal Site Reliability Engineer
About the Role and Job Responsibilities Define and evolve the organization's Site Reliability Engineering strategy, vision, and technical roadmap in partnership with engineering leadership and business stakeholders. Guide the lifecycle of services from conception to end-of-life across the organization, including establishing service design review frameworks, capacity planning methodologies, and production readiness standards that scale across teams. Establish and drive adoption of organization-wide standards and best practices related to system architecture, service delivery, reliability, automation, and operational excellence. Influence technical decisions across multiple teams and business units. Build and operate shared platforms, tooling, and frameworks that enable service and engineering teams across the organization to achieve high availability and improve incident detection and response. Drive system performance, availability, and efficiency improvements across the organization through architectural guidance, automation strategy, process refinement, and deep analysis of operational patterns in incident reviews. Partner closely with engineering and product leadership across the organization to establish reliability standards, shape technical direction, and deliver reliable services at scale. Champion a culture of operational excellence by treating operational challenges such as software engineering problems and mentoring teams on reducing toil at scale. Lead, mentor, and develop Site Reliability Engineering talent across multiple teams, establishing best practices and growing SRE capability within the organization. Partner with executive stakeholders, product leadership, and business teams to align reliability investments with business priorities and drive technical strategy that supports business outcomes.
Qualifications 12+ years of hands-on experience in software engineering, systems engineering, and/or cloud-based environments. 10+ years of experience working with public cloud platforms (e.g., GCP, AWS, or Azure), including designing and operating large-scale systems. 5+ years of experience designing, operating, and maintaining applications and/or systems infrastructure in large-scale, customer-facing production environments. Demonstrated experience architecting and influencing large-scale distributed systems and infrastructure solutions. Proven track record of leading cross-functional technical initiatives and mentoring engineering teams. Demonstrated understanding of observability best practices, including metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing. Experience coding in one or more higher-level programming languages (e.g., Python, Java, C# or C++). Strong working knowledge of Linux systems, including troubleshooting, performance analysis, and scripting in production environments. Experience with GitHub Actions and modern CI/CD practices. Exceptional communication and collaboration skills, with demonstrated ability to influence across teams, lead technical discussions with executives, and mentor engineers.
Preferred Qualifications Experience defining and communicating technical vision and strategy across large organizations. Deep expertise in distributed system design and architecture Hands-on experience with cloud-native applications and containerization technologies (e.g., Kubernetes, containers). Experience designing and managing infrastructure-as-code and configuration management strategies across organizations (e.g., Terraform, Ansible). Experience operating and optimizing production workloads at scale, including cost optimization and performance tuning. Solid grounding in at least three of the following areas: Computer Science fundamentals, Cloud Architecture, Security or Network Design Experience designing observability strategies and building metrics pipeline, operational dashboards and alerts using observability tools such as Splunk or Grafana. Experience partnering with business and product teams on technology decisions and translating technical recommendations into business value. Experience driving operational change and building consensus around new technical practices and standards
Company Overview:
UKG is the Workforce Operating Platform that puts workforce understanding to work. With the world's largest collection of workforce insights, and people-first AI, our ability to reveal unseen ways to build trust, amplify productivity, and empower talent, is unmatched. It's this expertise that equips our customers with the intelligence to solve any challenge in any industry — because great organizations know their workforce is their competitive edge. Learn more at ukg.com.
Equal Opportunity Employer
UKG is an equal opportunity employer. We evaluate qualified applicants without regard to race, color, disability, religion, sex, age, national origin, veteran status, genetic information, and other legally protected categories.
View The EEO Know Your Rights poster
UKG participates in E-Verify. View the E-Verify posters here.
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Disability Accommodation in the Application and Interview Process
For individuals with disabilities that need additional assistance at any point in the application and interview process, please email UKGCareers@ukg.com.
The pay range for this position is $184,300 to $264,950. The actual base pay offered may vary depending on skills, experience, job-related knowledge and work location. In addition to base pay, employees may be eligible to participate in a performance-based bonus plan and to receive restricted stock unit awards as part of total compensation. Learn more about UKG’s benefits and rewards at https://www.ukg.com/about-us/careers/benefits
Similar roles
-
Site Reliability Engineer, Data Center Infrastructure
SpaceX Bastrop, Texas, United States
-
Sr. Site Reliability Engineer, Platform Infrastructure
SpaceX Bastrop, Texas, United States
-
Software Engineer, SRE and Production Engineering - DGX Cloud
NVIDIA Santa Clara, California, United States · $184K–$288K/yr
-
Production Site Reliability Engineer
Charles Schwab Inc. Omaha, Nebraska, United States · $120K–$155K/yr
-
Site Reliability Engineer (Associate, Experienced, or Senior)
Boeing Berkeley, Missouri, United States · $99K–$217K/yr
-
Site Reliability Engineer
Chevron Buenos Aires, Argentina