JPMorgan Chase & Co.

Site Reliability Engineer III

JPMorgan Chase & Co. Hyderabad, Telangana, India

Financial Services · 10,001+ employees

7 h ago Closes in 2d
sre Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and support scalable, resilient AI/ML data solutions while coordinating incident management and performing root cause analysis. Additionally, you will mentor team members and leverage AI capabilities to optimize operational workflows and reduce toil.

What they look for

Site Reliability Engineering Incident Management Root Cause Analysis Python PySpark Observability Grafana Dynatrace Prometheus Datadog Splunk System Design Disaster Recovery AWS Databricks Snowflake

Requirements

Candidates must have 3+ years of applied experience in site reliability engineering and proficiency in incident resolution and observability tools. Strong knowledge of Python or PySpark and experience with cloud platforms like AWS, Databricks, or Snowflake are required.

Full description

Join a dynamic team where your expertise in site reliability engineering will shape the future of AI/ML data platforms. Unlock opportunities for growth and impact as you help build resilient, market-leading solutions.

As a Site Reliability Engineer III at JPMorgan Chase within the AI/ML Data Platforms team, you will play a pivotal role in developing scalable and resilient data solutions. You will engage in root cause analysis, production changes, and strategic initiatives that drive operational excellence. Your experience will help mentor team members and foster collaboration across global teams. Together, we create innovative solutions that support the firm’s mission and community.

Job responsibilities

  • Build and support scalable, resilient AI/ML data solutions
  • Coordinate incident management coverage for effective application issue resolution
  • Collaborate with cross-functional teams to perform root cause analysis and implement production changes
  • Develop and support AI/ML solutions for troubleshooting and incident resolution
  • Mentor and guide team members to drive strategic change
  • Manage budgetary considerations and staffing challenges
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Partner with colleagues across global teams to deliver impactful results

Required qualifications, capabilities and skills

  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience
  • Proficient in site reliability culture and principles, with familiarity in implementing site reliability within an application or platform
  • Proficiency in running production incident calls and managing incident resolution
  • Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
  • Strong understanding of SLI/SLO/SLA, Error Budgets and Proficiency in Python or PySpark for AI/ML modeling
  • Must be able to reduce toil by building new tools to automate repeated tasks
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Awareness of risk controls and compliance with departmental and company-wide standards
  • Ability to work collaboratively in teams and build meaningful relationships to achieve common goals

Preferred qualifications, capabilities and skills

  • 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies
  • AWS and Databricks certifications

Similar roles