Principal Site Reliability Engineer (Python, Java and Automation Specialization)
Oracle India
IT Services and IT Consulting · 10,001+ employees
About the role
The Principal Site Reliability Engineer designs and architects scalable infrastructure while collaborating with software teams to ensure system reliability. They are responsible for forecasting capacity needs, performing incident response, and identifying opportunities for automation.
What they look for
Requirements
Candidates should have 3 to 5+ years of experience in site reliability engineering or a related field. The role requires advanced knowledge of infrastructure maintenance, performance reporting, and the ability to conduct experiments with new tools.
Benefits
Full description
Job Responsibilities
- Own and improve the reliability, scalability, performance, and operational efficiency of critical OCI Compute services.
- Lead investigation and resolution of complex production incidents; drive mitigation, recovery, RCA, and follow-up improvements.
- Improve service KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures.
- Build automation and tooling to reduce recurring operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support major upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot complex distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Mentor team members and act as a technical resource for partner teams.
- Participate in a 12x7 on-call rotation and lead response during customer-impacting incidents.
Mandatory Skills
- Total Experience of 7+ Years
- 5+ years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available distributed systems in production.
- Strong programming or scripting skills in Python, Java, Go, or similar languages.
- Strong hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience defining or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident-management, troubleshooting, RCA, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on complex technical issues and collaborate across engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or large-scale rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience mentoring engineers or leading technical initiatives across teams.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions
- Does the candidate have 5+ years of SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
- Has the candidate owned or significantly improved a critical production service or infrastructure component?
- Can they lead a complex incident end-to-end: investigation, mitigation, recovery, RCA, and follow-up actions?
- Do they have strong hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
- Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
- Have they built or improved automation, CI/CD pipelines, deployment validation, operational tooling, or remediation workflows?
- Do they have experience with monitoring, alerting, dashboards, logs, metrics, tracing, and SLO/KPI-driven reliability improvement?
- Can they work independently on ambiguous technical problems, collaborate across engineering teams, mentor peers, and participate in a 12x7 on-call rotation?
Role Details
Field
Requirement
Role
Site Reliability Engineer 4 (IC4)
Experience
5+ Years
Location
Bangalore Only
Work Mode
Hybrid - 3 days from Office
Primary Skills
Site Reliability Engineering, Production Engineering, Linux, OCI/Cloud Infrastructure, Distributed Systems, Python/Java/Go, Automation, CI/CD, Monitoring and Observability, Incident Management, RCA, SLOs/SLIs/KPIs, Performance Tuning, Capacity Planning, AIOps, Deployment Validation, On-Call Operations, Security Vulnerability Management, Security, Vulnerability
35% Ops, 65% - Automation and Coding
Qualifications
Career Level - IC4
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
Similar roles
-
Staff Site Reliability Engineer - Federal
ServiceNow San Diego, California, United States · $150K–$262K/yr
-
Senior Site Reliability Engineer
Synapse Health $134K–$184K/yr
-
Senior Site Reliability Engineer
Mirantis Hyderabad, Telangana, India
-
Site Reliability Engineering Lead
Experian Hyderabad, Telangana, India
-
Senior Site Reliability Engineer
Fivetran Novi Sad, Vojvodina, Serbia
-
Senior Site Reliability Engineer
Veeam Software pune, Maharashtra, India