Senior Site Reliability Engineer, Production Engineering
Jobgether India
Internet Marketplace Platforms · 11-50 employees
About the role
You will support and maintain large-scale production Kubernetes services within a global 24/7 reliability operations environment. This includes automating operational tasks, managing incidents, and collaborating with cross-functional teams to ensure service availability and reliability.
What they look for
Requirements
Candidates must have 7+ years of experience administering large-scale production Kubernetes environments and a bachelor's degree in a relevant field. Strong expertise in Linux systems administration, networking, and troubleshooting complex infrastructure issues is required.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer, Production Engineering based in India.
This is a senior production engineering opportunity focused on keeping large-scale cloud and infrastructure services highly available, reliable, and secure. You’ll support production Kubernetes services within a global 24/7 reliability operations environment, with a strong emphasis on automation and reducing manual operational work. The role combines deep systems expertise with incident management, observability, cluster administration, and reliability engineering. You’ll work closely with Site Reliability Engineering, security, DevOps, development teams, and service owners to prevent and resolve complex production issues. Your work will directly contribute to service availability, customer experience, and operational resilience at significant scale. The position is well suited to an engineer who enjoys solving complex infrastructure problems and working with advanced cluster technologies.
\n
Accountabilities:
• Support production Kubernetes services as part of a global 24/7 production engineering operation, including flexibility to work split-weekend shifts.
• Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure while protecting service availability, integrity, reliability, and SLAs.
• Automate operational processes and continuously identify opportunities to reduce manual tasks and improve engineering efficiency.
• Use monitoring, observability, alerts, and alarms to proactively detect, prevent, investigate, and respond to production incidents.
• Analyze logs, metrics, system behavior, and infrastructure signals to troubleshoot complex issues and determine root causes.
• Lead incident management calls, coordinating timely detection, escalation, investigation, and resolution of critical production issues.
• Engage subject matter experts, service owners, and cross-functional engineering teams to resolve complex incidents efficiently.
• Develop and improve monitoring, alerting, and reliability mechanisms in collaboration with development teams.
• Perform systems administration and security monitoring across large-scale infrastructure environments.
• Apply deep knowledge of Linux, networking, Kubernetes, and cluster infrastructure to maintain reliable production services.
• Contribute to the architecture, deployment, and ongoing improvement of Kubernetes environments operating at significant scale.
• Continuously evaluate emerging infrastructure and high-performance computing technologies and identify opportunities for innovation.
Requirements:
• 7+ years of demonstrated experience administering large-scale production Kubernetes environments within high-availability Internet, cloud, or data-center environments, with strong on-premises experience preferred.
• Bachelor's degree in Computer Science, Engineering, Mathematics, or a related discipline, or equivalent professional experience.
• Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management.
• Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
• Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core Linux networking.
• Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure.
• Experience with CI/CD technologies and tools such as Jenkins and ArgoCD.
• Scripting or programming experience in Python, Golang, or Rust is preferred but not mandatory.
• Strong understanding of observability, incident management, reliability engineering, and production operations.
• Excellent analytical and troubleshooting skills, with the ability to work effectively under pressure during complex incidents.
• Strong communication and interpersonal skills, including the ability to clearly present technical information and influence cross-functional stakeholders.
• Ability to learn new technologies quickly and adapt to evolving infrastructure environments.
• Experience architecting, building, and deploying Kubernetes environments at large scale is highly valuable.
• Passion for innovation and advanced high-performance cluster technologies is an advantage.
Benefits:
• Full-time opportunity with a remote working option in India.
• Opportunity to work on large-scale production Kubernetes and infrastructure environments.
• Exposure to advanced cloud, bare-metal, GPU/DPU, and high-performance computing technologies.
• Opportunity to work alongside SRE, DevOps, security, development, and other specialized engineering teams.
• Significant technical ownership across reliability, automation, observability, incident response, and infrastructure operations.
• Opportunity to solve complex engineering challenges at global scale.
• Continuous exposure to emerging technologies and opportunities to develop advanced infrastructure expertise.
• 24/7 production engineering environment offering substantial experience in incident management and high-availability operations.
• Compensation, healthcare, leave, and other employment benefits are provided according to the applicable employment package and location-specific terms.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Associate Site Reliability Engineer
Culture Amp Melbourne, Victoria, Australia · A$90K/yr
-
Senior/Lead Site Reliability Engineer
NetEase Games Singapore, Singapore
-
Senior Site Reliability (DevOps) Engineer (INPD)
Rakuten Singapore, Singapore
-
Staff Site Reliability Expert
Lightspeed Commerce, Inc. Auckland, Auckland, New Zealand
-
Site Reliability Engineer
MyFitnessPal United States · $120K–$160K/yr
-
Senior Platform Engineer (SRE)
KKCompany Technologies Taipei, Taiwan