Site Reliability Engineer III
JPMorgan Chase & Co. Hyderabad, Telangana, India
Financial Services · 10,001+ employees
About the role
You will build and support scalable, resilient AI/ML data solutions while coordinating incident management and performing root cause analysis. Additionally, you will mentor team members and leverage AI capabilities to optimize operational workflows and reduce toil.
What they look for
Requirements
Candidates must have 3+ years of applied experience in site reliability engineering and proficiency in incident resolution and observability tools. Strong knowledge of Python or PySpark and experience with cloud platforms like AWS, Databricks, or Snowflake are required.
Full description
Join a dynamic team where your expertise in site reliability engineering will shape the future of AI/ML data platforms. Unlock opportunities for growth and impact as you help build resilient, market-leading solutions.
As a Site Reliability Engineer III at JPMorgan Chase within the AI/ML Data Platforms team, you will play a pivotal role in developing scalable and resilient data solutions. You will engage in root cause analysis, production changes, and strategic initiatives that drive operational excellence. Your experience will help mentor team members and foster collaboration across global teams. Together, we create innovative solutions that support the firm’s mission and community.
Job responsibilities
- Build and support scalable, resilient AI/ML data solutions
- Coordinate incident management coverage for effective application issue resolution
- Collaborate with cross-functional teams to perform root cause analysis and implement production changes
- Develop and support AI/ML solutions for troubleshooting and incident resolution
- Mentor and guide team members to drive strategic change
- Manage budgetary considerations and staffing challenges
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
- Partner with colleagues across global teams to deliver impactful results
Required qualifications, capabilities and skills
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Proficient in site reliability culture and principles, with familiarity in implementing site reliability within an application or platform
- Proficiency in running production incident calls and managing incident resolution
- Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
- Strong understanding of SLI/SLO/SLA, Error Budgets and Proficiency in Python or PySpark for AI/ML modeling
- Must be able to reduce toil by building new tools to automate repeated tasks
- Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
- Awareness of risk controls and compliance with departmental and company-wide standards
- Ability to work collaboratively in teams and build meaningful relationships to achieve common goals
Preferred qualifications, capabilities and skills
- 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies
- AWS and Databricks certifications
Similar roles
-
Staff Site Reliability Engineer, Core Networking
Google Dublin, Leinster, Ireland · €150K–€153K/yr
-
Sr. Staff Software Engineer – SRE, Release & Test Platforms
ServiceNow Dublin, Leinster, Ireland
-
Senior Site Reliability Engineer
Planet Berlin, Germany · €65K–€96K/yr
-
Site Reliability Engineer / Devops — Retail Engineering
Apple Shanghai, Shanghai, China
-
Software Engineer, Site Reliability Engineering, Colossus SRE
Google Dublin, Leinster, Ireland · €122K–€125K/yr
-
Site Reliability Engineer II
Akamai Krakow, Lesser Poland Voivodeship, Poland