EY

Site Reliability & Resilience -Senior manager

EY Tada, Andhra Pradesh, India

Professional Services · 10,001+ employees

Jul 22 Closes in 6d
sre Principal (10+ yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior Manager will act as a technical advisor to strategize and implement SRE transformation roadmaps across the enterprise. They are responsible for defining SLA/SLO/SLI metrics, designing observability solutions, and optimizing IT infrastructure costs through automation.

What they look for

Site Reliability Engineering Java J2EE CI/CD Terraform Kubernetes Observability Dynatrace Splunk Python Microservices Cloud Computing Performance Tuning Linux FinOps SLA Management

Requirements

Candidates must have over 15 years of experience in software product engineering and deep expertise in SRE principles. Proficiency in Java, cloud technologies, CI/CD tools, and performance monitoring is required for this role.

Full description

At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.

Site Reliability Engineering (SRE) Sr. Consultant (SM)

Description

  • 15 Years of experience in Site Reliability Engineering (SRE) is a modern way of delivering IT Operations by imbibing Software engineering principles in Service Delivery to reduce IT Risk to business, improve business resilience, attain predictability & reliability, optimize cost of IT Infra and Ops
  • An SRE Architect / Consultant will be a technical advisor to strategize the transformation roadmap to modernize IT delivery with SRE principles, frameworks and levers – in a nutshell, setting up SRE into Enterprise IT
  • They will also implement the SRE Roadmap and govern SRE solutions across the enterprise / line-of-businesses
  • They will be able to assess SRE Maturity of an IT Organization and provide strategy and roadmap to achieve higher maturity levels

Responsibilities

  • Defining SLA/SLO/SLI for a product / service
  • Engineering in resilient design and implementation practices into solutions as they go through the product life cycle
  • Designing & implementing Observability Solutions to track, report, and measure SLA adherence
  • Engineering out manual effort (Toil) through the development of automated processes and services (e.g., Automated Management of Systems, CI/CD improvements)
  • Optimize Cost of IT Infra & Operations - FinOps
  • Review, Analysis and Improvement of deployed products with respect to product architecture and inter-service dependencies - Simplification

Typical Skills and Background

  • 15+ years of experience in software product engineering principles, processes and systems
  • Hands-on experience in Java / J2EE, one of web server (Apache Tomcat or IBM HTTP Server), one of the application servers (Tomcat/WebSphere), and any major RDBMS like Oracle
  • Hands-on experience in at least one CI-CD (Azure DevOps, GitLab CI/CD, Jenkins) and IaC tools (Terraform, AWS CloudFormation, Ansible etc.)
  • Experience in at least one cloud technology (AWS/Azure/GCP etc. and Docker, Pivotal, Kubernetes, OpenShift etc.) and its reliability tools (Azure AppInsight, CloudWatch, Azure Monitor etc.)
  • Experience in Observability - APM tools (Dynatrace, AppDynamics etc.), metrics / log consolidation (Splunk) and ELK Stack
  • Defining NFRs and SLA/SLO/SLI agreement for a product / platform / services
  • Knowledge on queuing models used, thread pools, request servicing processes etc.
  • Experience in Linux (RHEL) operating system performance monitoring parameters and their interpretation, commands used for monitoring
  • Experience in Web Services, SOA, ESB (DataPower), RESTFul
  • Knowledge of application design patterns, J2EE application architectures, Microservices, Spring boot & Cloud native architectures
  • Proficiency in Java runtimes, Core Java, Garbage collection, JVM parameters tuning
  • Experience in performance tuning on Application Servers (Tomcat/WAS)
  • Experience in trouble shooting Performance / Scalability / Availability issues
  • Thread dump, heap dump generation & analysis
  • Knowledge on Query tuning and database architecture
  • Knowledge at least one automation scripting language like Python
  • Mastery of collaborative software development using Git, Jira, Confluence etc.
  • AI/ML & Data Analytics knowledge and experience is a desirable

EY | Building a better working world

EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets.

Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate.

Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today.

Similar roles