Jalasoft

Site Reliability Engineer - Azure, Observability and Scripting

Jalasoft · Peru

Software Development · 1,001-5,000 employees

4 h ago
Remote Senior (5-10 yrs) Full-time Peru
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will ensure the reliability, scalability, and performance of cloud-native platforms on Microsoft Azure and Kubernetes. Responsibilities include managing observability, automating infrastructure, and handling incident response and disaster recovery.

What they look for

Azure Kubernetes Observability Terraform Bicep Python PowerShell Bash Prometheus Grafana Azure DevOps Incident response Infrastructure as code Disaster recovery KQL Log Analytics

Requirements

Candidates must have 6+ years of experience, including 3+ years operating Kubernetes in production environments. Proficiency in Azure, monitoring tools, infrastructure as code, and scripting languages is required.

Benefits

Remote work 13 floating holidays 15 vacation days per year

Full description

Jalasoft is seeking a Site Reliability Engineer (SRE) to join our team and help ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes. The ideal candidate is passionate about observability, automation, and operational excellence, with experience improving system availability, monitoring, incident response, and infrastructure automation in production environments.

Required Years of Experience:

  • 6+ years of experience
  • 3+ years of experience operating Kubernetes in production

Must Haves:

  • Site reliability engineering or production operations for Kubernetes workloads at scale.
  • Azure Monitor, Log Analytics and KQL, including workspace design, data collection rules and retention strategy.
  • Prometheus and Grafana: metrics and exporters, recording and alerting rules, and dashboard design.
  • Definition and implementation of service level indicators, objectives and error budgets.
  • Alerting and incident response design, including runbook authoring and on-call practice.
  • Backup, restore and disaster recovery design and testing, including validation against RPO and RTO targets.
  • Kubernetes operations: workload troubleshooting, resource management and cluster upgrades.
  • Ability to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines.
  • Scripting in Python, PowerShell or Bash.
  • Professional working English.

Nice-to-Have:

  • Azure Managed Prometheus and Azure Managed Grafana.
  • OpenTelemetry instrumentation and distributed tracing.
  • Azure Backup, Azure Site Recovery, and snapshot-based recovery of virtual machines.
  • Chaos engineering or structured game day practice.
  • Database-layer observability, particularly for Oracle.
  • Cost and capacity management for AKS estates.
  • Incident management tooling and postmortem practice.
  • Certification: CKA, AZ-400 or equivalent.
  • Remote work
  • 13 floating holiday
  • 15 vacation days per year completed
  • Good working environment

Every qualified candidate who meets the requirements outlined in the job description will be considered in this hiring process without distinction.

Furthermore, Jalasoft is an equal opportunity employer. We wholeheartedly embrace our responsibility to make employment decisions without regard to race, age, marital or social status, national origin, disability, sex, gender identity or expression, or any other characteristic or group of candidates or employees unrelated to their qualifications and suitability for the position. Our management is committed to upholding this policy with respect.