Site Reliability Engineer 2
Oracle Bengaluru, Karnataka, India
IT Services and IT Consulting · 10,001+ employees
About the role
The Site Reliability Engineer will support Oracle Analytics services by performing incident response, troubleshooting complex service issues, and driving mitigation efforts. They will also build and maintain operational tooling, automation, and monitoring systems to improve service reliability and scalability.
What they look for
Requirements
Candidates must have a BS or MS in Computer Science or Engineering with experience supporting cloud infrastructure and large-scale distributed applications. Strong proficiency in Linux, networking fundamentals, and scripting languages like Python or Bash is required.
Benefits
Full description
Site Reliability Engineer part of Oracle Analytics Service Excellence (OASE) SRE team partnering with Oracle Analytics development teams to improve the reliability, availability, performance, operational support and maturity of Oracle Analytics Cloud services.
The ideal candidate is a hands-on Site Reliability Engineer with strong analytical skills and the ability to read, understand, investigate, and safely troubleshoot existing application code. This role requires diagnosing complex service issues across distributed systems, automation, infrastructure, and enterprise applications—using logs, telemetry, code investigation, and operational data to drive issues from detection through mitigation and prevention.
OASE develops tools, technologies, processes, and data-driven operating practices that improve service uptime, reduce time to mitigation, and enable scalable cloud operations. The team builds and operates internal services, automation, dashboards, and reporting capabilities that support Oracle Analytics customers, engineering teams, partners, and business growth. The ideal candidate enjoys working in an agile, customer-focused environment. This role is centered on improving uptime through proactive monitoring, incident response, code-level troubleshooting, automation, capacity analysis, operational tooling, patching, remediation, and continuous service improvement.
Responsibilities
Responsibilities
- Perform SRE activities supporting Oracle Analytics customers, engineering teams, and release cycles in both pre-production and production environments.
- Participate in a follow-the-sun model providing 24x7 operational support for Oracle Analytics services.
- Respond to incidents, troubleshoot complex service issues, drive mitigation to completion, and contribute to root-cause analysis and post-incident actions.
- Read, understand, and troubleshoot existing application and service code to diagnose complex production issues, identify root causes, and support safe remediation.
- Develop deep expertise in Oracle Analytics services to prevent regressions, resolve customer issues effectively, and reduce recurring incidents.
- Build, maintain, and improve operational tooling, automation, dashboards, monitoring, runbooks, and knowledge-base documentation.
- Build and maintain AI-assisted and automation tools as needed to improve SRE workflows, including incident investigation, monitoring, remediation, operational reporting, and knowledge management.
- Develop scripts and services for monitoring, telemetry collection, capacity analysis, patching, remediation, and operational reporting.
- Analyze service health, workload patterns, CPU, memory, query activity, and capacity trends to identify reliability and scaling risks.
- Execute interim patches, hot fixes, upgrades, and maintenance activities with operational excellence.
- Partner with Development, Support, Product Management, and other engineering teams to investigate and resolve service failures and outages.
- Improve CI/CD processes, deployment practices, operational readiness, and release supportability.
- Share operational knowledge, and continuously improve team processes.
- Follow established security, compliance, change-management, and operational procedures.
Required Qualifications
- BS or MS in Computer Science, Engineering, or equivalent practical experience.
- Experience supporting cloud infrastructure, networking, applications, services, tools, and operational processes.
- Strong understanding of networking and TCP/IP fundamentals, including DNS, HTTP/HTTPS, TLS, load balancing, and service connectivity.
- Linux/Unix system-administration experience, including troubleshooting processes, memory, CPU, filesystem, and network issues.
- Experience developing, operating, or supporting cloud services and large-scale distributed applications in production.
- Demonstrated ability to troubleshoot complex technical issues methodically, including investigation of existing applications and code.
- Experience creating and maintaining technical documentation, runbooks, knowledge articles, and operational guides.
- Experience working in agile development and operational environments.
- Strong written and verbal communication skills, including the ability to work effectively with remote global teams.
- Ability to work independently, manage competing priorities, and participate in on-call, after-hours maintenance, and weekend support as needed.
Preferred Qualifications
- Experience with Oracle Analytics Cloud, Oracle Analytics Server, OBIS, BI Publisher, Oracle Database, Autonomous Database, MySQL, SQL Server, or NoSQL technologies.
- Two to four years of experience operating large-scale, customer-facing web applications or cloud services.
- Experience with OCI, AWS, Azure, or GCP compute, storage, networking, monitoring, and operational tooling.
- Programming and scripting experience with Python, Bash, JavaScript/Node.js, Ansible, and related technologies; Java experience is a plus.
- Ability to read, understand, troubleshoot, and safely modify existing enterprise application code.
- Familiarity with AI-assisted development tools, such as Codex and Claude Code, for software development, automation, investigation, and documentation.
- Experience with CI/CD and infrastructure automation tools such as Ansible, Puppet, Chef, Git, and deployment pipelines.
- Experience with cloud-native applications, containers, Kubernetes, microservices, and independently scalable services.
- Experience with REST APIs, service integrations, and automation workflows.
- Experience using Jira and Confluence for incident management, issue tracking, operational documentation, and collaboration.
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
Similar roles
-
Staff Site Reliability Engineer
KEV Group Toronto, Ontario, Canada · $150K–$180K/yr
-
Site Reliability Engineer - Data Platform
IMC Amsterdam, North Holland, Netherlands
-
Staff Software Engineer, Site Reliability Engineering, Vertex AI
Google Warsaw, Masovian Voivodeship, Poland · PLN 480K–PLN 492K/yr
-
Site Reliability Engineering Manager
Conifers.ai Tel-Aviv, Tel-Aviv District, Israel
-
Senior Site Reliability Engineer
Precisely International Jobs Bielsko-Biała, Silesian Voivodeship, Poland
-
Network SRE
JPMorgan Chase & Co. Buenos Aires, Argentina