Senior Site Reliability / Gitops Engineer
Jobgether India
Internet Marketplace Platforms · 11-50 employees
About the role
You will drive automation and GitOps practices while acting as an embedded technical lead to design and architect reusable infrastructure services. You are responsible for maintaining core services, troubleshooting complex infrastructure issues, and collaborating with global teams to ensure system reliability and performance.
What they look for
Requirements
The role requires a strong background in modern hosting architectures, Infrastructure as Code, and professional experience with Python and Kubernetes. Candidates must possess a product-oriented mindset, solid Linux knowledge, and excellent English communication skills to work effectively in a distributed environment.
Benefits
Full description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability / Gitops Engineer based in India.
Join a globally distributed Site Reliability Engineering team responsible for operating critical IT services at significant scale. You will drive an automation-first approach across private and public cloud environments, using Infrastructure as Code and GitOps practices to improve reliability and efficiency. The role combines hands-on engineering, technical leadership, architecture, and operational responsibility across complex infrastructure. You will help design reusable services, automate software operations, and strengthen infrastructure practices through modern engineering methods. Your work will directly influence the reliability, scalability, performance, and observability of production systems used across a global organization. You will collaborate with engineers, architects, operations teams, and support specialists while contributing feedback and improvements to open-source technologies. This is an opportunity to take ownership of challenging infrastructure initiatives while mentoring others and shaping next-generation SRE and automation practices.
\n
Accountabilities:
- Drive the development of automation and GitOps practices within the team while acting as an embedded technical lead.
- Collaborate closely with the infrastructure architecture function to align technical solutions with broader architecture objectives.
- Design and architect infrastructure services that can be delivered as reusable products across the organization.
- Develop and strengthen Infrastructure as Code practices by continuously improving automation, processes, consistency, and reusability.
- Automate software operations across private and public clouds while accounting for the complexity and operational requirements of distributed systems.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliability and continuity.
- Troubleshoot complex infrastructure issues, support capacity planning, investigate performance challenges, and improve system resilience.
- Implement and maintain observability, monitoring, and alerting solutions using technologies such as Prometheus, Grafana, and Elasticsearch.
- Collaborate with globally distributed engineering, operations, and support teams to resolve technical challenges and deliver reliable services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive or manual operational tasks.
- Share technical knowledge, experience, and best practices through design sessions, mentorship, collaborative implementation, and team development.
- Take final responsibility for resolving time-critical escalations and ensuring appropriate technical follow-through.
Requirements:
- Strong understanding of modern hosting architectures and an automation-first approach based on Infrastructure as Code across private and public cloud environments.
- A product-oriented mindset with an interest in building reusable infrastructure products rather than one-off solutions.
- Professional experience with Python development, including work on large or complex projects.
- Hands-on experience with Kubernetes or other container orchestration technologies.
- Proven experience managing and deploying cloud infrastructure through code and automation.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from distributed storage such as Ceph to database-backed systems.
- Hands-on experience administering enterprise Linux servers.
- Strong knowledge of cloud computing concepts, architectures, and technologies.
- Bachelor’s degree or higher, preferably in Computer Science, Engineering, or a related technical discipline.
- Strong English communication skills across email, chat, video calls, voice communication, and in-person collaboration.
- Ability to troubleshoot problems across the technology stack, from the Linux kernel through application and web layers, while knowing when to seek input from others.
- Flexibility, curiosity, and the ability to learn new technologies and approaches quickly.
- Ability to adapt to fast-changing technical environments and remain focused on delivering reliable outcomes.
- Experience working effectively within globally distributed teams.
- Passion for open-source technologies, with familiarity with Ubuntu or Debian considered valuable.
Benefits:
- Compensation shaped according to geographic location, experience, and performance, with regular compensation reviews.
- Performance-driven annual bonus or commission in addition to base compensation.
- Distributed work environment with twice-yearly in-person team sprints.
- USD 2,000 annual personal learning and development budget.
- Recognition and performance rewards.
- Annual holiday leave.
- Maternity and paternity leave.
- Team Member Assistance Program and Wellness Platform.
- Opportunities to travel internationally and meet colleagues in new locations.
- Priority Pass and travel upgrades for long-haul company events.
- Opportunity to work on large-scale cloud infrastructure, SRE, GitOps, Infrastructure as Code, automation, observability, and open-source technologies.
- Dedicated development time for larger technical initiatives and automation projects.
- Opportunities to provide technical mentorship and influence infrastructure engineering practices.
\nHow Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
Similar roles
-
Site Reliability Engineer III
JPMorgan Chase & Co. Jersey City, New Jersey, United States · $138K–$185K/yr
-
Site Reliability Engineer (Manufacturing Infrastructure)
SpaceX Bastrop, Texas, United States
-
IC3 - Infra Engineer - SRE
Spin Careers Ciudad de México, Mexico
-
Engineering Manager, Site Reliability
Booking Holdings Bangalore South, Karnataka, India
-
Senior Site Reliability Engineer
ServiceTitan California, United States · $138K–$221K/yr
-
Site Reliability Engineer 2
Oracle United Kingdom