Procter & Gamble

Site Reliability Engineer

Procter & Gamble · Taguig, Metro Manila, Philippines

Manufacturing · 10,001+ employees

14 h ago
Mid (2-5 yrs) Full-time Philippines
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Site Reliability Engineer will ensure the high availability and reliability of digital IT products by implementing robust monitoring and observability solutions. They will also lead initiatives to automate incident response, optimize infrastructure, and drive continuous improvement across IT systems.

What they look for

Site Reliability Engineering Monitoring Observability Automation Linux Unix Cloud Platforms Terraform Python C# Networking Protocols Load Balancing DNS Management Docker Kubernetes SQL

Requirements

Candidates must hold a Bachelor's degree in a related field such as Engineering, IT, or Computer Science with up to 5 years of experience. Proficiency in system administration, automation scripting, cloud platforms, and containerization technologies is required.

Full description

Job Location

Taguig City

Job Description

Information Technology (IT) at Procter & Gamble is where business, innovation and technology integrate to build a competitive advantage for P&G. Our mission is clear -- you deliver IT to help P&G win with consumers. 

 

Do you love implementing continuous improvement in IT solutions to drive efficiency and agility in meeting constantly evolving business needs? Then this job might be for you! 

 

As a Site Reliability Engineer, you will be instrumental in ensuring the high availability and reliability of our digital IT products in P&G.  Your primary focus will be on enhancing system performance through faster detection, response, and resolution of issues, while also implementing strategies to prevent recurrence and reduce operational toil. You will use robust Observability and Monitoring tools, automate incident response systems, and optimize IT architecture to create a resilient and reliable infrastructure. 

This is a Managerial position. Being a manager at P&G involves leading teams and / or end-to-end processes, managing P&G resources, and driving business results. Managers are responsible for overseeing various aspects of the business, including strategy, operations, and team performance. They play a crucial role in ensuring that P&G's brands continue to grow and succeed in the market. Managers at P&G are expected to have strong leadership skills, a growth mindset, and the ability to make data-driven decisions roles lead and initiatives, significantly impacting business results through independent judgment and minimal guidance.

Responsibilities: 

  • Implement and lead comprehensive monitoring solutions and tools to provide real-time insights into system performance, enabling proactive incident detection and ensuring accurate, actionable alerts for prompt responses. 
  • Continuously refine monitoring strategies and develop automation scripts to address recurring issues, enhancing system visibility, resource optimization, and overall efficiency. 
  • Establish and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to improve service quality and reliability, 
  • Collect and share data and insights from observability tools to drive continuous improvement initiatives. 
  • Work closely with Software Engineers, Product Teams, and Infrastructure Teams to develop and implement initiatives that enhance IT reliability. 
  • Engage with customers to understand their needs and difficulties regarding Observability and Monitoring tools, providing exceptional support in all interactions, including communications, updates, and feedback. 
  • Stay updated on industry trends and effective strategies in Site Reliability Engineering while continuously enhancing technical skills in system architecture, automation, cloud technologies, and operational processes. 

Job Qualifications

Candidates must demonstrate strong leadership in the application of technical expertise to drive business results. 

We are looking for candidates who possess the following core qualities: 

  • A Bachelor's degree in related field such as Engineering, Information Technology and Computer Science discipline, and up to 5 years experience at most. 
  • Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana) 
  • Knowledge and familiarity in system administration, including Linux/Unix environments, cloud platforms (Azure or GCP preferred, but AWS is acceptable) 
  • Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform) 
  • Proficiency in at least one programming language (e.g., Python, C#) and a background in scripting for automation tasks 
  • Understanding of networking protocols, network infrastructures, load balancing, and DNS management 
  • Familiarity with containerization and Orchestration Technologies (e.g., Docker, Kubernetes) 
  • Familiarity with databases and proficiency in writing SQL queries 
  • Understanding of best practices in security and experience with implementing secure systems 
  • Knowledge of incident response methodologies, root cause analysis, and implementing preventive measures (ITIL and/or SRE) 
  • Familiarity with ticketing systems and task management (preferably ServiceNow) 
  • Problem-solving skills with ability to analyze complex issues and devise effective solutions 
  • Learning agility as there will be new topics to learn and new spaces to understand 
  • Communication and collaboration skills to work effectively with multi-functional teams, partners, and customers 
  • Teamwork and interpersonal skills, with an ability to build relationships and work effectively in a collaborative environment 
  • Operational excellence / execution skills as the work requires discipline 

Job Schedule

Full time

Job Number

R000156530

Job Segmentation

Entry Level