Jobgether

Sr. Site Reliability Engineer (Azure, IaC, Distributed Systems)

Jobgether India

Internet Marketplace Platforms · 11-50 employees

13 h ago
Remote sre Mid (2-5 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will be responsible for automating operational processes, strengthening monitoring capabilities, and supporting resilient infrastructure across development, staging, and production environments. Additionally, you will investigate production incidents, perform root cause analysis, and implement improvements to enhance system performance and reliability.

What they look for

Azure Infrastructure as Code Distributed Systems Observability CI/CD Docker Kubernetes Terraform Ansible CloudWatch Splunk Prometheus Grafana OpenTelemetry Generative AI Incident Management

Requirements

Candidates must have a bachelor's degree with 2+ years of experience or 5+ years of professional experience in SRE or DevOps roles. Proficiency in cloud platforms, observability tools, Infrastructure as Code, and containerization technologies is required.

Benefits

Fully remote work model Full-time position Exposure to large-scale technology environments Professional growth opportunities Inclusive workplace Equal-opportunity employment

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr. Site Reliability Engineer (Azure, IaC, Distributed Systems) based in India.

This is an opportunity to contribute to the reliability, observability, and scalability of critical technology environments in a global organization.The role focuses on automating operational processes, strengthening monitoring capabilities, and supporting resilient infrastructure across development, staging, and production.You will work with cloud platforms, infrastructure-as-code, CI/CD pipelines, containers, and a broad observability stack.The position offers exposure to modern SRE and DevOps practices while collaborating with technical teams to improve system performance and operational efficiency.You will help investigate production incidents, identify root causes, and implement improvements that reduce recurring issues.The role also encourages the use of automation and emerging technologies, including Generative AI, to make engineering workflows more effective.This is a remote position with the opportunity to work on technically challenging infrastructure and reliability initiatives.

\n

Accountabilities:

  • Support the deployment, configuration, and maintenance of monitoring and logging solutions across development, staging, and production environments.
  • Maintain and optimize observability technologies, including Splunk, ClickHouse, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, and CloudWatch.
  • Automate repetitive operational activities and identify opportunities to improve workflows, efficiency, and system integration.
  • Contribute to the setup and maintenance of CI/CD pipelines supporting automated build, testing, and deployment processes.
  • Support cloud infrastructure management across AWS and GCP, with a focus on availability, reliability, and security.
  • Use Infrastructure as Code tools such as Terraform, Ansible, and CloudFormation to configure and manage environments.
  • Support the implementation and administration of containerization technologies, including Docker and Kubernetes.
  • Monitor system performance, identify potential issues, and escalate complex incidents appropriately.
  • Participate in production incident troubleshooting, root cause analysis, and the implementation of corrective and preventive actions.
  • Provide first-level support for infrastructure and deployment issues while collaborating with other technical teams on more complex problems.
  • Create and maintain clear documentation covering infrastructure, operational processes, configurations, and procedures.
  • Apply DevOps and SRE best practices while contributing to continuous improvements in platform reliability and operational maturity.
  • Leverage emerging technologies, including Generative AI tools, to support day-to-day engineering activities and improve productivity.

Requirements:

  • Bachelor’s degree with 2+ years of relevant experience, or 5+ years of relevant professional experience without the degree requirement.
  • Hands-on experience supporting monitoring and logging tool deployment and configuration in production environments.
  • Practical experience with observability platforms such as Splunk, Grafana, Prometheus, OpenTelemetry, Fluent Bit, Elasticsearch, OpenSearch, ClickHouse, or CloudWatch.
  • Experience supporting cloud infrastructure, particularly AWS and/or GCP, with an understanding of availability, security, and operational reliability.
  • Experience with Infrastructure as Code technologies such as Terraform, Ansible, and CloudFormation.
  • Experience building or maintaining CI/CD pipelines for automated software delivery and deployment.
  • Familiarity with Docker and Kubernetes and their use in modern infrastructure environments.
  • Experience monitoring system performance and supporting production incident management, troubleshooting, and root cause analysis.
  • Understanding of DevOps and Site Reliability Engineering principles and a willingness to continuously develop technical expertise.
  • Strong analytical and problem-solving skills, with the ability to investigate issues systematically and identify opportunities for automation.
  • Ability to work collaboratively with engineering and infrastructure teams while managing priorities in a dynamic environment.
  • Strong documentation and communication skills, with attention to operational detail and process consistency.
  • Digital fluency and willingness to incorporate modern tools, including Generative AI solutions, into everyday engineering workflows.

Benefits:

  • Fully remote work model in India.
  • Full-time position with a standard Monday–Friday work schedule.
  • Opportunity to work with modern cloud, observability, automation, containerization, and Infrastructure as Code technologies.
  • Exposure to large-scale technology environments and distributed systems.
  • Opportunities to develop expertise in DevOps, Site Reliability Engineering, cloud infrastructure, and automation.
  • Collaborative environment with opportunities to work across technical teams and global initiatives.
  • Opportunity to use emerging technologies, including Generative AI, to improve engineering productivity and operational processes.
  • Professional growth through hands-on experience with complex infrastructure and reliability challenges.
  • Inclusive workplace and equal-opportunity employment environment.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles