Oracle

Site Reliability Engineer 2

Oracle Bengaluru, Karnataka, India

IT Services and IT Consulting · 10,001+ employees

5 d ago
sre Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will be responsible for improving the reliability, scalability, and performance of OCI Compute services through automation and operational excellence. This includes investigating complex production incidents, performing root cause analysis, and collaborating with development teams to maintain service-level objectives.

What they look for

Site Reliability Engineering Python Go Linux Cloud infrastructure Networking Compute Storage Monitoring Alerting Distributed systems Capacity planning Performance tuning Incident management Automation CI/CD

Requirements

Candidates must have 2-4 years of experience in SRE or systems engineering with strong programming skills in Python or Go. Proficiency in Linux, cloud infrastructure, and distributed systems troubleshooting is required to support production environments.

Benefits

Medical insurance Life insurance Retirement options Volunteer programs

Full description

At Oracle Cloud Infrastructure (OCI), we are building the future of cloud for enterprises with the agility of a startup and the scale of a global enterprise leader. Compute is one of OCI’s foundational organizations, responsible for delivering the core infrastructure powering Virtual Machines (VMs) and Bare Metal (BM) services.

As an OCI Site Reliability Engineer (SRE), you will work closely with development and product teams in a shared full-stack ownership model across multiple services and technology domains. You will develop deep expertise in service architecture, dependencies, configurations, and operational behavior across large-scale production environments.

You will be responsible for improving the reliability, scalability, performance, and operational efficiency of OCI Compute services. The role includes handling critical customer incidents, supporting deployments, performing validation and operational testing, troubleshooting complex infrastructure issues, conducting root cause analysis (RCA), and driving service reliability improvements.

You will act as a key escalation point for complex production issues, leveraging strong knowledge of distributed systems, service topology, and infrastructure dependencies to identify mitigations and restore service health while partnering with development teams to meet SLA commitments.

The role also involves leveraging AIOps and intelligent automation to enhance monitoring, anomaly detection, event correlation, predictive alerting, RCA, and remediation workflows. Using observability platforms, telemetry analytics, and automation frameworks, you will help reduce operational toil, improve incident response, and enhance overall service reliability.

This is an opportunity to combine deep technical expertise with operational excellence to solve complex cloud infrastructure challenges at massive scale within Oracle’s next-generation cloud platform.

Responsibilities

Job Responsibilities

  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, and production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Mandatory Skills

  • 2-4 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, security vulnerability management or production rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
  • Familiarity with security, compliance, and access-control practices in production environments.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Similar roles