Engineering Manager, SRE
Jobgether Netherlands · $75K–$170K/yr
Internet Marketplace Platforms · 11-50 employees
About the role
Lead and develop a Site Reliability Engineering team while maintaining hands-on involvement in technical direction and infrastructure challenges. You will be responsible for maturing reliability practices, including SLOs and observability, while balancing operational excellence with long-term engineering initiatives.
What they look for
Requirements
Requires proven experience leading SRE or infrastructure teams with deep technical expertise in Kubernetes, AWS, and cloud operations. Candidates must demonstrate strong leadership skills, experience in regulated environments, and the ability to work effectively in a distributed, asynchronous setting.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Netherlands.
This is a hands-on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform. You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges. The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering. You will shape how the team balances operational excellence, incident response, reliability improvements, and longer-term engineering initiatives. A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization. You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership. The environment is fully remote and asynchronous, offering significant autonomy in a fast-growing, globally distributed organization.
\n
Accountabilities
- Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring.
- Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus.
- Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly.
- Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed.
- Design and maintain effective support rotations and on-call processes while strengthening incident response practices.
- Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform.
- Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post-incident improvements.
- Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations.
- Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership.
- Remain hands-on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalate.
- Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early.
- Continuously improve team health, collaboration, conflict resolution, and retrospective practices.
Requirements:
- Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team, with direct responsibility for performance and career development.
- Strong hands-on background in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review designs, challenge implementation decisions, and contribute during production incidents.
- Production experience with Kubernetes and AWS at meaningful scale, including the operational realities of running cloud infrastructure.
- Hands-on experience building, enabling, or scaling AI infrastructure and working with AI-related engineering workloads.
- Strong understanding of observability principles and practices, infrastructure as code with Terraform, and CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins.
- Experience with Docker, shell scripting, and production infrastructure operations.
- Proven ownership of reliability practices including incident response, on-call operations, SLOs, error budgets, and turning incidents into lasting engineering improvements.
- Experience working in regulated environments, with an understanding of infrastructure controls, compliance, and security requirements.
- Exceptional prioritization skills, particularly when operational workloads compete with project commitments.
- Excellent written communication and documentation skills, with the ability to lead effectively in a highly distributed and asynchronous environment.
- Strong relationship-building, collaboration, conflict-resolution, and stakeholder-management capabilities.
- A coaching-oriented leadership style, with evidence of developing engineers both technically and professionally.
- Strong judgment, accountability, adaptability, curiosity, and commitment to high-quality execution.
- Nice-to-have experience with Elixir, Java, Clojure, Node.js, Python, or another backend programming language.
- Additional desirable experience includes OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora performance optimization, connection pool management, query tuning, Linux systems administration, security, FinOps, and cloud cost management.
- Experience growing an engineering team from a small base and establishing a strong hiring bar is advantageous.
- Ability to work effectively across global teams and time zones.
Benefits:
- Annual salary range of USD $75,450–$169,700, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
- Fully remote, work-from-anywhere environment.
- Flexible working hours within an asynchronous work culture.
- Flexible paid time off.
- 16 weeks of paid parental leave.
- Budget for coworking spaces, learning, and wellness activities, including gym memberships.
- Mental health support services.
- Stock options.
- Home office budget and IT equipment.
- Global exposure through collaboration with colleagues across multiple continents.
- Opportunities to travel internationally and meet colleagues at company events.
- A high-autonomy environment where employees are encouraged to organize their schedules around their lives and personal commitments.
- Opportunity to influence the maturity of reliability engineering practices while working on complex infrastructure and platform challenges.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Staff Site Reliability Engineer, Core Networking
Google Dublin, Leinster, Ireland · €150K–€153K/yr
-
Sr. Staff Software Engineer – SRE, Release & Test Platforms
ServiceNow Dublin, Leinster, Ireland
-
Senior Site Reliability Engineer
Planet Berlin, Germany · €65K–€96K/yr
-
Site Reliability Engineer / Devops — Retail Engineering
Apple Shanghai, Shanghai, China
-
Software Engineer, Site Reliability Engineering, Colossus SRE
Google Dublin, Leinster, Ireland · €122K–€125K/yr
-
Site Reliability Engineer II
Akamai Krakow, Lesser Poland Voivodeship, Poland