Jobgether

Site Reliability Engineer (SRE) - Engineering Productivity

Jobgether India

Internet Marketplace Platforms · 11-50 employees

9 h ago
Remote sre Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, build, and operate scalable, secure, and highly available infrastructure systems while automating operational workflows. Collaborate with development teams to resolve technical bottlenecks, improve observability, and manage incident response procedures.

What they look for

Site Reliability Engineering Go Python Shell Scripting Linux Unix Infrastructure as Code Automation Observability Cloud Computing Docker Kubernetes Prometheus Grafana CI/CD Troubleshooting

Requirements

Requires a Bachelor’s or Master’s degree in a technical field and approximately 5+ years of relevant experience. Candidates must possess strong skills in Linux/Unix administration, automation scripting, and infrastructure-as-code practices.

Benefits

Large-scale infrastructure exposure Modern technology stack access High autonomy and ownership Collaborative engineering culture Continuous learning opportunities Global engineering environment

Full description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer (SRE) - Engineering Productivity based in India.

As a Site Reliability Engineer, you will help build and operate the infrastructure that powers software engineering and product development at scale. You will design secure, resilient, and highly available systems across hybrid cloud environments. The role combines software engineering, infrastructure, automation, observability, and production operations. You will work closely with development teams to remove technical bottlenecks and improve the overall developer experience. You will have ownership of critical systems, from deployment and monitoring through incident response and continuous improvement. The environment is highly engineering-focused, collaborative, and designed to encourage learning and experimentation with modern technologies. This is an opportunity to make a direct impact on engineering productivity while working with large-scale, business-critical infrastructure.

\n

Accountabilities:

  • Design, build, deploy, and operate critical production systems with a strong focus on scalability, reliability, observability, performance, and security.
  • Develop automation that reduces operational toil and improves the efficiency of production systems and engineering workflows.
  • Monitor infrastructure and services proactively, improve alerting, and implement automated responses where appropriate.
  • Create, maintain, and continuously improve incident response procedures and operational runbooks.
  • Build and deploy new systems incrementally using staged rollout practices to minimize operational risk.
  • Investigate and resolve platform and infrastructure issues, supporting software engineering teams with technical triage and troubleshooting.
  • Collaborate with third-party vendors when required to diagnose and resolve infrastructure or platform-related issues.
  • Write post-incident reviews and implement corrective measures to prevent recurring incidents.
  • Plan and communicate production maintenance activities and maintenance windows.
  • Partner with product development teams to identify infrastructure-related bottlenecks and design effective solutions.
  • Implement fault-tolerance, performance improvements, and scaling strategies to increase system availability and resilience.
  • Research and adopt infrastructure and platform best practices to maintain secure, scalable, and fault-tolerant environments.
  • Develop a strong understanding of open-source and industry-standard systems to improve troubleshooting, diagnosis, and issue resolution.
  • Support and enhance the overall developer experience across internal platforms and services.

Requirements:

  • Hold a Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or possess equivalent professional experience, with approximately 5+ years of relevant experience.
  • Have working knowledge of Go, Python, and/or shell scripting, with the ability to develop medium-complexity automation workflows.
  • Demonstrate strong Linux or UNIX administration and debugging capabilities.
  • Have hands-on experience operating software systems, infrastructure, or complex applications at scale.
  • Have experience with server provisioning, particularly from storage and networking perspectives.
  • Possess strong software troubleshooting, analytical, and problem-solving skills.
  • Have practical experience with infrastructure-as-code and automated infrastructure management.
  • Be comfortable working across multiple technologies and adapting quickly to new tools and platforms.
  • Experience with databases such as MariaDB, PostgreSQL, MongoDB, or similar technologies is desirable.
  • Experience with Docker, virtualization, or related technologies such as KVM, QEMU, or Kata Containers is a plus.
  • Experience managing monitoring and observability platforms such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos is beneficial.
  • Experience managing Elasticsearch clusters, Artifactory, container registries, or similar infrastructure platforms is advantageous.
  • Familiarity with CI/CD technologies such as ArgoCD or Spinnaker is desirable.
  • Experience with version-control systems such as Perforce or Gerrit is a plus.
  • Experience with infrastructure-as-code frameworks such as Ansible is beneficial.
  • Experience managing large Java applications or storage infrastructure such as NAS, SAN, or Ceph is an advantage.
  • Demonstrate strong communication and collaboration skills, with the ability to work effectively with software engineers, infrastructure teams, and external technology partners.

Benefits:

  • Opportunity to work on large-scale, business-critical infrastructure and engineering productivity systems.
  • Exposure to a broad range of modern cloud, automation, observability, CI/CD, database, storage, and infrastructure technologies.
  • Significant ownership and autonomy over engineering projects and technical solutions.
  • Collaborative environment with opportunities to work across different technical domains and development teams.
  • Hybrid cloud engineering experience across scalable and fault-tolerant systems.
  • Opportunities for continuous learning, experimentation, and adoption of infrastructure best practices.
  • Engineering-focused culture that emphasizes technical excellence, automation, quality, and innovation.
  • Opportunity to contribute to systems and tools that directly improve the productivity and experience of software development teams.
  • Access to a globally distributed engineering environment with opportunities for cross-functional and international collaboration.

\nHow Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Similar roles