AI-Centric Site Reliability Engineer
Capgemini Hyderabad, Telangana, India
IT Services and IT Consulting · 10,001+ employees
About the role
Implement and operate AI-centric SRE practices to drive reliability, observability, and automated remediation across enterprise infrastructure. Design self-healing workflows and leverage AI-driven tools to accelerate root cause analysis and improve service resiliency.
What they look for
Requirements
Requires 8+ years of IT experience with at least 5 years in production support or SRE roles. Candidates must possess strong expertise in the Dynatrace observability platform and AI-powered operational tools.
Benefits
Full description
Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world.
Your Role
- Implement and operate AI-centric SRE practices across enterprise applications and infrastructure.
- Leverage Dynatrace OneAgent, Smartscape, Distributed Tracing, Real User Monitoring (RUM), Synthetic Monitoring, and Davis AI for proactive monitoring and anomaly detection.
- Utilize Davis AI for root cause analysis, event correlation, predictive problem management, and automated remediation recommendations.
- Design and implement self-healing and autonomous operational workflows.
- Drive observability strategy across application, infrastructure, middleware, databases, and cloud platforms.
Your Profile
- 8+ Years of IT Experience with 5+ Years in Production Support, Site Reliability Engineering, Application Support, or Platform Operations.
- Experienced AI-Centric Site Reliability Engineer (SRE) to drive reliability, observability, automation, and intelligent operations across a diverse enterprise technology landscape.
- Strong expertise in Dynatrace Observability Platform, Davis AI, Model Context Protocol (MCP) Servers, Claude Code, AI-powered Self-Service Agents, and modern ITSM and Agile practices.
- This role requires hands-on experience in managing complex production environments, leveraging AI-driven observability and automation to reduce incidents, improve service resiliency, accelerate root cause analysis, and enhance customer experience.
- The candidate should have strong functional understanding of applications spanning distributed systems, databases, middleware, mainframe technologies, infrastructure platforms, content management systems, and insurance platforms.
What you'll love about working here
- You can shape your career with us. We offer a range of career paths and internal opportunities within Capgemini group. You will also get personalized career guidance from our leaders.
- You will get comprehensive wellness benefits including health checks, telemedicine, insurance with top-ups, elder care, partner coverage or new parent support via flexible work.
- At Capgemini, you can work on cutting-edge projects in tech and engineering with industry leaders or create solutions to overcome societal and environmental challenges.
Capgemini is an AI-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organizations and make it real with AI, technology and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end-to-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations. The Group reported 2024 global revenues of €22.1 billion. Make it real | www.capgemini.com
Similar roles
-
Site Reliability Engineer
Kaluza Bristol, England, United Kingdom · £56K–£84K/yr
-
Staff Site Reliability Engineer, AQI Data SRE
Google Pittsburgh, Pennsylvania, United States · $207K–$300K/yr
-
Software Engineering Manager II, Site Reliability Engineering, Data Intelligence
Google San Jose, California, United States · $207K–$300K/yr
-
Copy of Senior Site Reliability Engineer (EMEA)
Reap Poland, Line Islands, Kiribati
-
Senior Site Reliability Engineer (EMEA)
Reap Poland, Line Islands, Kiribati
-
Tech Lead SRE(f/h/n)
Hublo Paris, Ile-de-France, France · €85K–€95K/yr