Staff Site Reliability Engineer, Government
Jobgether United States
Internet Marketplace Platforms · 11-50 employees
About the role
Lead the design and continuous improvement of highly reliable, scalable, and resilient systems for government-focused environments. Partner with cross-functional teams to establish engineering standards, resolve complex production challenges, and mentor staff to ensure operational excellence.
What they look for
Requirements
Requires extensive professional experience in site reliability or infrastructure engineering with a proven track record of staff-level technical leadership. Candidates must possess strong expertise in distributed systems, cloud infrastructure, and observability, along with the ability to influence architecture across multiple teams.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer, Government based in the United States.
This is a senior engineering opportunity focused on building and maintaining reliable infrastructure for government-focused technology environments. The role sits within a research and development organization and offers the opportunity to influence reliability practices at significant technical scale. You’ll apply deep expertise in site reliability engineering to improve availability, resilience, observability, and operational performance. As a Staff-level engineer, you’ll contribute beyond individual systems by shaping architecture, engineering standards, and operational strategy. You’ll work across engineering teams to solve complex reliability challenges and build scalable approaches that support mission-critical workloads. The position is well suited to an experienced infrastructure leader who combines strong technical judgment with a systems-level perspective. This is a fully remote role for professionals based in the United States, supporting a technically sophisticated government-focused environment.
\n
Accountabilities:
- Lead the design, implementation, and continuous improvement of highly reliable, scalable, and resilient systems supporting government-focused environments.
- Establish and influence site reliability engineering practices, standards, and architectural approaches across engineering teams.
- Identify and address systemic reliability, availability, scalability, performance, and operational risks.
- Develop automation and engineering solutions that reduce manual operational work and improve the consistency of production environments.
- Strengthen observability, monitoring, alerting, incident response, and service health practices.
- Partner with software, infrastructure, security, and other technical teams to resolve complex production and reliability challenges.
- Provide technical leadership on architecture and engineering decisions with long-term reliability and operational excellence in mind.
- Mentor and guide engineers while helping raise engineering standards and reliability practices across the organization.
- Contribute to incident management, root-cause analysis, and the implementation of sustainable corrective actions.
- Help establish and evolve reliability objectives, operational processes, and engineering best practices for critical systems.
- Influence technical strategy across projects and teams through strong systems thinking, collaboration, and technical judgment.
Requirements:
- Extensive professional experience in site reliability engineering, infrastructure engineering, platform engineering, DevOps, or a closely related discipline.
- Staff-level technical leadership experience, with a demonstrated ability to influence architecture, engineering practices, and technical strategy across multiple teams.
- Strong understanding of distributed systems, cloud infrastructure, production operations, automation, and system reliability principles.
- Experience designing and operating highly available and scalable production systems.
- Strong expertise in observability, monitoring, alerting, incident response, and troubleshooting complex systems.
- Proven ability to identify systemic technical problems and develop durable, scalable solutions rather than short-term fixes.
- Strong programming or scripting capabilities for automation and infrastructure tooling.
- Excellent communication and collaboration skills, with the ability to work effectively across engineering and technical disciplines.
- Strong technical judgment and the ability to make pragmatic decisions in complex, high-impact environments.
- Experience mentoring engineers and influencing teams without relying solely on formal authority.
- Ability to operate effectively in a remote environment and independently manage complex technical initiatives.
- Experience working with government, regulated, security-sensitive, or mission-critical environments is highly valuable.
Benefits:
- Remote position within the United States.
- Opportunity to work on technically complex, reliability-critical systems supporting government-focused environments.
- Staff-level scope with significant influence over engineering practices, architecture, and reliability strategy.
- Opportunity to collaborate with experienced engineering and research & development teams.
- Professional environment focused on technical innovation, scalability, and operational excellence.
- Additional salary, healthcare, retirement, paid time off, and other benefits were not specified in the provided job description.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Sr Technical Consultant - Azure Cloud, DevOps, Sql & Site Reliability Engineering(SRE)
Blue Yonder Bengaluru, Karnataka, India
-
Site Reliability Engineer I
PagerDuty Atlanta, Georgia, United States · $98K–$148K/yr
-
Site Reliability Engineer II
PagerDuty Atlanta, Georgia, United States · $113K–$172K/yr
-
Site Reliability Engineer, Enterprise Technology Services
Apple Shanghai, Shanghai, China
-
AWS Cloud Site Reliability Engineer
UnitedHealth Group Basking Ridge, New Jersey, United States · $73K–$130K/yr
-
Site Reliability Engineer
Bay Systems Consulting Inc. Berkeley, California, United States