Sr. SRE Engineer - Dynatrace, Python, AIOPS
Chubb · Hosakote taluk, Karnataka, India
Insurance · 1,001-5,000 employees
About the role
The Senior SRE Engineer will design and operate automation and self-healing systems to improve production reliability and reduce manual toil. They will also lead incident response, define SLI/SLO metrics, and mentor junior team members on engineering best practices.
What they look for
Requirements
Candidates must have at least 5 years of experience in software engineering or SRE roles with hands-on proficiency in Python or Go. A bachelor degree in a technical discipline is required, along with experience in observability platforms like Dynatrace and cloud infrastructure.
Benefits
Full description
About Chubb
Chubb is a world leader in insurance. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. The company is defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb employs approximately 40,000 people worldwide. Additional information can be found at: www.chubb.com.
About Chubb India
At Chubb India, we are on an exciting journey of digital transformation driven by a commitment to engineering excellence and analytics. We are proud to share that we have been officially certified as a Great Place to Work® for the third consecutive year, a reflection of the culture at Chubb where we believe in fostering an environment where everyone can thrive, innovate, and grow
With a team of over 2500 talented professionals, we encourage a start-up mindset that promotes collaboration, diverse perspectives, and a solution-driven attitude. We are dedicated to building expertise in engineering, analytics, and automation, empowering our teams to excel in a dynamic digital landscape.
We offer an environment where you will be part of an organization that is dedicated to solving real-world challenges in the insurance industry. Together, we will work to shape the future through innovation and continuous learning.
Position Details
- Job Title: Sr. SRE Engineer - Dynatrace, Python, AIOPS
- Function/Department: Technology
- Location: Hyderabad/Bangalore
- Employment Type: Full Time
- Reports To: TADEPALLI, MADHURI
Position Summary:
- Act as a Senior, hands-on engineer within the SRE team — designing, building, and operating automation and self-healing systems that reduce toil and strengthen production reliability.
- Take ownership of SLI/SLO instrumentation and error budget tracking for an assigned portfolio of services, partnering with development teams to drive measurable reliability improvements.
- Serve as a senior on-call escalation point and incident responder, driving root cause analysis and durable hardening fixes for critical production issues.
- Champion observability best practice through hands-on Dynatrace instrumentation, dashboard design, and alert tuning across owned services.
- Mentor junior and mid-level SREs, raise engineering standards across the team, and contribute technical input to the team's automation and reliability roadmap.
Major Duties and Responsibilities
- Technical Leadership & Mentorship
- Act as a technical role model within the SRE team; set the standard for engineering quality, automation-first thinking, and reliability best practice.
- Mentor and coach junior and mid-level SREs on troubleshooting techniques, automation approaches, and reliability engineering principles.
- Review peers' designs, runbooks, and automation scripts; provide constructive feedback that raises the bar for engineering quality.
- Contribute to technical interviews and skills assessment of SRE candidates when requested.
- Support the SRE Lead in shaping the team's reliability roadmap by proposing, prototyping, and validating improvements.
- Incident Response & On-Call
- Participate in the 24/7 follow-the-sun on-call rotation as a senior escalation point (L2/L3); take ownership of complex, high-severity incidents.
- Act as incident commander or technical lead for P1/P2 incidents when rostered; drive diagnosis, mitigation, and recovery under pressure.
- Author and maintain blameless postmortems for incidents you own; track corrective actions through to closure.
- Build and continuously improve runbooks and playbooks for services you support; automate manual steps wherever feasible — a runbook unchanged after 90 days is a toil backlog item.
- Partner with Problem Management on root cause investigations; implement permanent fixes rather than workarounds.
- Reliability Engineering & SLO Ownership
- Define and instrument SLI/SLO measurements for assigned services; monitor error budget burn and flag risk to release velocity.
- Build automation and self-healing tooling (Python, Go, or equivalent) that reduces manual toil and improves MTTD/MTTR for owned services.
- Conduct capacity planning and performance analysis for owned services; recommend scaling and architecture improvements.
- Participate in Production Readiness Reviews (PRR): assess new services against reliability criteria and provide sign-off recommendations.
- Conduct Failure Mode and Effects Analysis (FMEA) for services you support; prioritise and implement hardening fixes.
- Embed reliability practices into the SDLC for owned services — operability reviews, readiness checklists, and reliability testing as standard gates.
- Chaos Engineering & Resilience
- Design and execute chaos engineering experiments and GameDays for owned services; document findings and drive remediation.
- Implement fault injection, load testing, and synthetic failure scenarios as part of CI/CD pipelines.
- Apply resilience patterns — circuit breakers, graceful degradation, bulkhead isolation — to harden owned services.
- Maintain and update the resilience scorecard for your service portfolio.
- Observability & Dynatrace
- Instrument applications and infrastructure using Dynatrace (OneAgent, OpenTelemetry) to achieve full-stack observability for owned services.
- Build and maintain SLO dashboards, health scorecards, and Davis AI alerting rules for owned services.
- Tune alert thresholds and reduce noise; continuously improve the signal-to-noise ratio for your service portfolio.
- Contribute to team-wide observability standards: instrumentation guidelines, tagging taxonomy, and log retention policy.
- AI-Augmented Operations
- Build and maintain automation scripts and tooling that integrate with AIOps and LLM-based operational tools — incident summarisation, runbook generation, and knowledge retrieval.
- Use AI-assisted triage and anomaly detection tools to accelerate diagnosis; provide feedback to improve model accuracy.
- Pilot new AI/ML-based reliability tooling under guidance from the SRE Lead; document outcomes and recommendations.
- Collaboration & Continuous Improvement
- Partner with development and platform teams to influence service design for reliability and operability at design time, not post-deployment.
- Contribute to quarterly SRE KPI reporting: SLO attainment, incident trends, and toil metrics for owned services.
- Support knowledge transfer for project-to-support transitions; validate production readiness before go-live.
- Stay current with SRE best practice (Google SRE principles, DORA metrics) and bring emerging techniques into the team.