Staff Site Reliability Engineer
Jobgether · United States
Internet Marketplace Platforms · 11-50 employees
About the role
The Staff Site Reliability Engineer will lead technical strategy for reliability, observability, and platform engineering to ensure scalable and resilient production systems. They will mentor engineering teams, drive automation initiatives, and define standards for incident response and operational excellence.
What they look for
Requirements
Candidates must have 12+ years of experience in software or infrastructure engineering, with at least 6 years specifically in SRE environments. Expert-level knowledge of cloud platforms, distributed systems, container orchestration, and programming languages like Python or Go is required.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in United States.
We are seeking a Staff Site Reliability Engineer to serve as a senior technical authority responsible for shaping reliability strategy and production excellence.This role goes beyond system maintenance, focusing on building scalable platforms, improving engineering practices, and driving long-term operational improvements.You will influence how distributed systems, cloud infrastructure, and AI-driven reliability practices evolve across the organization.The position combines deep technical expertise with strategic leadership, mentorship, and cross-functional collaboration.You will help define standards for observability, incident response, automation, and platform engineering at scale.This is an opportunity to make a significant impact on the reliability of mission-critical systems while helping teams build safer and more efficient technology.
\n
Accountabilities: The Staff Site Reliability Engineer will lead the technical direction of reliability engineering initiatives, ensuring production systems remain scalable, secure, and resilient. This role requires advanced expertise in distributed systems, cloud platforms, automation, and operational excellence while serving as a trusted advisor to engineering teams.
- Define and execute technical strategy across observability, alerting, platform infrastructure, and reliability engineering practices.
- Lead the evolution of scalable, secure, and efficient cloud platforms and distributed systems.
- Establish and improve reliability frameworks including SLIs, SLOs, error budgets, capacity planning, and operational readiness processes.
- Drive automation initiatives using software engineering practices, Infrastructure as Code, and platform capabilities to reduce operational complexity.
- Lead investigations into complex production incidents and transform learnings into permanent engineering improvements.
- Build self-service platform capabilities that improve developer productivity, system safety, and operational ownership.
- Shape observability strategies, monitoring solutions, anomaly detection approaches, and AI-powered reliability improvements.
- Partner with engineering leadership and technical stakeholders to make critical architectural and reliability decisions.
- Mentor engineers across different experience levels and elevate reliability practices throughout the organization.
- Communicate technical risks, opportunities, and recommendations clearly to engineering, product, and executive audiences.
- Support the development of sustainable on-call strategies, tooling, and production support practices.
Requirements:
The ideal candidate is a highly experienced reliability engineering professional with strong software engineering skills and a proven ability to lead complex technical initiatives. Success in this role requires deep expertise in cloud infrastructure, distributed systems, automation, and technical leadership.
- 12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering.
- 6+ years of experience working in SRE environments and 3+ years leading complex cross-functional technical initiatives.
- Expert-level knowledge of observability, incident response, capacity planning, automation, and reliability engineering practices.
- Advanced experience with container orchestration platforms, especially Kubernetes.
- Experience with observability platforms such as New Relic, Datadog, or equivalent solutions.
- Strong programming skills in Python, Go, Bash, or another general-purpose language.
- Experience building production tooling, automation frameworks, and platform capabilities.
- Strong understanding of distributed systems, cloud infrastructure, and scalable architecture.
- Proven ability to mentor engineers and influence technical direction across teams.
- Excellent communication skills with the ability to explain complex technical concepts to diverse audiences.
- Experience working in regulated environments such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is strongly preferred.
- Knowledge of AI/ML applications in reliability engineering, including AIOps, anomaly detection, automated remediation, and resource optimization.
Benefits:
- Competitive and fair compensation package.
- Medical, dental, and vision insurance for eligible employees.
- Maternity and paternity leave benefits.
- Short-term and long-term disability coverage.
- Opportunity to work within a rapidly growing technology environment.
- Ability to learn from and collaborate with an experienced leadership team.
- Access to company-provided equipment and branded merchandise.
- Opportunity to influence large-scale reliability practices and modern engineering approaches.
- Environment focused on innovation, technical excellence, and continuous learning.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1