NTT Data Singapore

Site Reliability Engineer - HM: Mukesh

NTT Data Singapore Singapore, Singapore

IT Services and IT Consulting · 10,001+ employees

8 h ago
sre Senior (5-10 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will ensure system health, reliability, and performance through automation and observability. They will collaborate with engineering teams to define SLIs, SLOs, and error budgets while driving initiatives to reduce operational toil.

What they look for

Site Reliability Engineering Observability Automation Fault Tolerance ITIL Infrastructure Deployment Bash Python Chaos Engineering SLI SLO Error Budgets MTTD MTTR Vendor Management Project Management

Requirements

Candidates must have a strong understanding of SRE and ITIL processes, along with proficiency in Bash or Python scripting for infrastructure deployment. Leadership, vendor management, and project management skills are required, with AWS certification considered a plus.

Full description

As a Site Reliability Engineer you will be filling a mission-critical role ensuring that our systems are healthy, monitored, automated, fault tolerant and designed to scale.

You will collaborate and work closely with engineering teams to continually improve our production services, facilitating fast delivery of new products, and reducing downtime.

Key Responsibilities:

  • Drive Site Reliability Engineering agenda to improve availability, reliability, and performance of services
  • Drive observability for our applications.
  • Drive optimise-operate initiative, example, reduction of operation toil
  • Work with application teams in setting up SLI, SLO and Error budget for their applications
  • Work with enterprise team in deploying SRE enablers/initiatives.

Requirements:

  • Have a good understanding of ITIL & SRE processes & practices
  • Have good leadership skills in working with application teams and service providers in defining infrastructure deployment plan, cutover/migration strategy and test plan.
  • Able to formulae and establish infrastructure deployment standards.
  • Good people management, vendor management and project management skills
  • Agile, AWS certification preferred
  • Able to create Bash/Python scripts for infra deployment
  • Must able to practice SRE & Chaos Engineering principles
  • Understands key SRE concepts such as Toil, SLI, SLO, Error Budgets, MTTD, MTTR, etc
  • Strong, committed, and reliable team player, able to take direction but also willing to contribute to discussions on design and strategy.
  • Possess strong interpersonal and communication skills to be able to deal with and form good relationships with other technology teams through day to day support and project work

Similar roles