About the role
The Site Reliability Engineer is responsible for ensuring the availability, performance, and scalability of enterprise platforms through monitoring and automation. They will lead incident management, define SLIs and SLOs, and implement Infrastructure as Code practices to maintain resilient systems.
What they look for
Requirements
Candidates should have 5 to 10 years of experience and proficiency in cloud platforms like Azure, AWS, or GCP. Strong technical skills in observability tools, container orchestration, and CI/CD pipeline management are required.
Full description
•
- Job Summary
- The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies
- Technical Skills
- Cloud Platforms: Azure
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
Responsibilities
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces) Define and track SLIs, SLOs, and SLAs to measure system performance Lead/support incident management, root cause analysis (RCA), and post-incident reviews Automate operational tasks and implement Infrastructure as Code (IaC) practices Support and improve CI/CD pipelines for stable and efficient releases 3. Skills & Competencies
Qualifications
- Technical Skills
- Cloud Platforms: AWS / Azure / GCP
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog Containers & Orchestration: Docker, Kubernetes CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps Infrastructure as Code: Terraform, Ansible, CloudFormation
At Zensar, we’re “experience-led everything”. We are committed to conceptualizing, designing, engineering, marketing, and managing digital solutions and experiences for over 130 leading enterprises. We are a company driven by a bold purpose: Together, we shape experiences for better futures. Whether for our clients, our people, or the world around us, this belief powers everything we do. At the heart of our culture is ONE with Client - a set of four core values that reflect who we are and how we work: One Zensar, Nurturing, Empowering, and Client Focus.
Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. Explore Life at Zensar and join us to Grow. Own. Achieve. Learn. to be the best version of yourself.
We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Similar roles
-
Staff Site Reliability Engineer
Anduril Industries Costa Mesa, California, United States · $191K–$253K/yr
-
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Elastic Poland · PLN 359K–PLN 465K/yr
-
Junior Site Reliability Engineer
Semarchy United States
-
Site Reliability Engineer
TradeStation Heredia, Costa Rica
-
Analista Júnior em Infraestrutura | SRE
XP Inc. São Paulo, São Paulo, Brazil
-
Analista Pleno em Infraestrutura | SRE
XP Inc. São Paulo, São Paulo, Brazil