Site Reliability Engineer - Azure, Observability and Scripting
Jalasoft · Peru
Software Development · 1,001-5,000 employees
About the role
The Site Reliability Engineer will ensure the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes. Responsibilities include managing observability, automating infrastructure, and handling incident response and disaster recovery.
What they look for
Requirements
Candidates must have 6+ years of experience, including 3+ years operating Kubernetes in production environments. Proficiency in Azure, monitoring tools, infrastructure as code, and scripting languages is required.
Benefits
Full description
Jalasoft is seeking a Site Reliability Engineer (SRE) to join our team and help ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes. The ideal candidate is passionate about observability, automation, and operational excellence, with experience improving system availability, monitoring, incident response, and infrastructure automation in production environments.
Required Years of Experience:
- 6+ years of experience
- 3+ years of experience operating Kubernetes in production
Must Haves:
- Site reliability engineering or production operations for Kubernetes workloads at scale.
- Azure Monitor, Log Analytics and KQL, including workspace design, data collection rules and retention strategy.
- Prometheus and Grafana: metrics and exporters, recording and alerting rules, and dashboard design.
- Definition and implementation of service level indicators, objectives and error budgets.
- Alerting and incident response design, including runbook authoring and on-call practice.
- Backup, restore and disaster recovery design and testing, including validation against RPO and RTO targets.
- Kubernetes operations: workload troubleshooting, resource management and cluster upgrades.
- Ability to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines.
- Scripting in Python, PowerShell or Bash.
- Professional working English.
Nice-to-Have:
- Azure Managed Prometheus and Azure Managed Grafana.
- OpenTelemetry instrumentation and distributed tracing.
- Azure Backup, Azure Site Recovery, and snapshot-based recovery of virtual machines.
- Chaos engineering or structured game day practice.
- Database-layer observability, particularly for Oracle.
- Cost and capacity management for AKS estates.
- Incident management tooling and postmortem practice.
- Certification: CKA, AZ-400 or equivalent.
- Remote work
- 13 floating holiday
- 15 vacation days per year completed
- Good working environment
Every qualified candidate who meets the requirements outlined in the job description will be considered in this hiring process without distinction.
Furthermore, Jalasoft is an equal opportunity employer. We wholeheartedly embrace our responsibility to make employment decisions without regard to race, age, marital or social status, national origin, disability, sex, gender identity or expression, or any other characteristic or group of candidates or employees unrelated to their qualifications and suitability for the position. Our management is committed to upholding this policy with respect.