Tech Lead Manager, Site Reliability
Jobgether United States · $146K–$183K/yr
Internet Marketplace Platforms · 11-50 employees
About the role
You will lead a team of Site Reliability Engineers while remaining hands-on with coding, architecture, and infrastructure modernization. The role involves driving system reliability, improving incident response, and fostering team development through coaching and performance management.
What they look for
Requirements
Candidates must have at least 4 years of professional software engineering experience and 3 years of platform engineering experience. You need strong technical leadership skills and the ability to act as the final decision-maker for distributed systems and modern infrastructure.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Tech Lead Manager, Site Reliability based in the United States.
This role combines hands-on software engineering, Site Reliability Engineering, and people leadership within a high-ownership technical team. You will lead a small team of 3–4 Site Reliability Engineers while remaining deeply involved in coding, architecture, infrastructure, and incident response. The position is designed for a technical leader who can own both engineering direction and team development without relying on a separate Tech Lead. You will drive reliability, platform adoption, infrastructure modernization, and improvements to developer self-service and delivery speed. Success will be measured through stronger system performance, faster incident detection and recovery, reduced operational toil, and higher-quality engineering practices. This is an opportunity to shape both the technology and the way a growing engineering team operates.
\n
Accountabilities:
- Serve as a hands-on contributor to Site Reliability and platform engineering work, including incident response, infrastructure automation, tooling, and technical improvements.
- Own technical debt and infrastructure modernization across the team's scope, prioritizing and resolving issues directly while balancing reliability with delivery needs.
- Maintain a deep technical understanding of the team's projects, including observability, disaster recovery, infrastructure, and platform initiatives, and actively challenge and improve technical approaches.
- Own team delivery outcomes and engineering quality, with accountability for reliability and performance metrics, incident response effectiveness, and adoption of platform tools.
- Guide technical direction through system design, architectural consultation, technical decision-making, and hands-on code reviews.
- Lead and develop a team of 3–4 Site Reliability Engineers, including career development, performance management, hiring, coaching, and overall team health.
- Partner with engineering leadership to evolve Agile, Scrum, or Kanban processes based on what works effectively for the team.
- Balance delivery urgency with appropriate engineering quality, reliability, and long-term maintainability.
- Model effective AI-assisted software development through hands-on use of AI tools and establish strong practices for the wider team.
- Participate in onboarding, codebase exploration, infrastructure work, planning, reviews, on-call activities, and incident response.
- Contribute directly to technical work within the first months while developing strong relationships with direct reports and gaining a detailed understanding of the team's systems and workflows.
- Join the on-call rotation and identify opportunities to improve response times, reduce repeated alerts, and strengthen incident management.
- Lead significant technical or architectural decisions, including improvements to observability, disaster recovery, and infrastructure.
- Identify and resolve technical or platform risks before they become delivery problems, including through disaster recovery and incident response exercises.
- Establish a sustainable balance between hands-on technical contribution and people leadership while becoming a trusted technical authority for Site Reliability.
Requirements:
- 4+ years of professional software engineering experience with strong, current hands-on coding experience.
- 3+ years of professional platform engineering experience incorporating DevOps and Site Reliability Engineering principles.
- Previous experience owning technical direction, people management, or both, with clear readiness to combine technical leadership and people leadership responsibilities.
- Ability to serve as the final technical decision-maker for a team without relying on a separate Tech Lead.
- Experience with distributed systems and modern infrastructure technologies such as AWS, Kubernetes/EKS, kops, HashiCorp Vault, Grafana, GitHub Actions, PostgreSQL, Kafka/MSK, Redis/Valkey, or comparable platforms.
- Experience with Azure or GCP environments is also welcome.
- Strong understanding of modern microservice architectures and engineering methodologies.
- Experience with Golang, TypeScript, and React is a plus.
- Experience with AI-assisted development tools and workflows, or a strong willingness and aptitude to adopt them.
- Demonstrated ability to make pragmatic technical trade-offs while working under real delivery and operational pressure.
- Experience with Agile or Scrum-based development processes, or willingness to help evolve team processes based on practical experience.
- Strong written and verbal communication skills, with the ability to collaborate effectively across technical and non-technical stakeholders.
- Strong coaching, mentoring, and people leadership capabilities.
- Ability to balance strategic technical thinking with direct execution and hands-on engineering contribution.
Benefits:
- National target base salary of $146,000–$183,000, depending on experience and skills.
- Participation in an annual bonus and stock option program.
- Fully flexible work model, with the option to work remotely, from the Pittsburgh office, or through a combination that suits your needs.
- 100% employer-paid employee health plan, including vision, dental, and supplemental coverage.
- Flexible Paid Time Off policy.
- Work-from-home stipend to help create a personalized and effective home office.
- 12 weeks of fully paid parental leave for all employees, plus short-term disability for birthing parents.
- 401(k) plan with employer matching.
- Opportunity to make an immediate impact on products used by millions of people.
- High-ownership environment where technical ideas and contributions can directly influence engineering practices and company outcomes.
- Remote interview process designed to provide a flexible and comfortable candidate experience.
- Employment is available to candidates legally authorized to work in the United States without current or future sponsorship.
- Hiring is currently limited to states where the organization already has employees; relocation assistance is not provided.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Technical Program Manager, Infrastructure SRE
Google Warsaw, Masovian Voivodeship, Poland · PLN 348K–PLN 356K/yr
-
Platform Engineer / SRE - Montpellier - H/F
Iliad - Free Montpellier, Occitania, France
-
Platform Engineer / SRE - Paris - H/F
Iliad - Free Paris, Ile-de-France, France
-
Site Reliability Engineer – Cloud Native Platform (OpenShift)
KPN Amersfoort, Utrecht, Netherlands · €71K–€110K/yr
-
Director Principal SRE
Barclays pune, Maharashtra, India
-
SRE Engineer
Nitka Technologies United States