Platform Site Reliability Engineer (SRE)
Broadridge Manila, Metro Manila, Philippines
Financial Services · 10,001+ employees
About the role
The Platform Site Reliability Engineer is responsible for ensuring the stability, scalability, and reliability of enterprise products through automation and proactive monitoring. They will collaborate with cross-functional teams to implement SRE best practices, manage system capacity, and conduct root cause analysis for technical incidents.
What they look for
Requirements
Candidates must hold a bachelor's degree in a relevant field and possess at least 5 years of experience in SRE or DevOps roles. Proficiency in cloud platforms, containerization, and automation tools is required, along with strong problem-solving skills and experience with observability systems.
Full description
At Broadridge, we've built a culture where the highest goal is to empower others to accomplish more. If you’re passionate about developing your career, while helping others along the way, come join the Broadridge team.
Role Overview
As a Platform Site Reliability Engineer (SRE), you will play a critical role in ensuring the stability, scalability, and reliability of our products and services. You will work closely with cross-functional teams to design, develop, and deploy solutions that enhance the performance and uptime of our applications.
The Platform SRE is part of the Enterprise Platform (EP) group and is responsible for supporting and running our standard platforms efficiently and effectively. You will be expected to collaborate closely with other functions within EP (DevOps/Cloud Platforms, Quality Engineering and Developer Experience) to provide robust, integrated and best-in-class solutions for our product engineering teams.
Key Responsibilities
- Implementing Site Reliability Engineering best practices, including error budgeting, service level objectives (SLOs), and monitoring and alerting systems
- Building automation tools and processes to improve the efficiency and reliability of running our products and standard platforms
- Performing capacity planning and system design to ensure that our systems can handle increasing traffic and load
- Troubleshooting complex technical issues and providing root cause analysis to prevent future incidents
- Participating in incident calls to respond to system outages and emergencies
- Collaborating with software developers to define and implement reliability requirements for new products/ applications/services
- Conducting post-mortem analyses to identity opportunities for improvement and prevent recurring issues
- Using data-based decision making to be proactive in the prevention of potential incidents and problems
- Supporting product development teams in the implementation of tools, processes, and practices to improve stability, reliability, and extensibility of their products
- Collaborate across the EP function to ensure that standard platforms are best-in-class
- Drive standard implementation of NFRs in new product development and own the "deep-dive" process to improve problematic application
- Overall management and governance of vulnerabilities and End-of-life within our products
As a senior member of the team, you will be responsible for:
- Overseeing initiatives and deliverables across the team
- Technical and design decisions made by the team.
- Coaching and mentoring members of the team
- Keeping your finger on the pulse: identifying and developing new ideas and initiatives
- Acting as an advocate for SRE across Enterprise Platform team and wider Broadridge community
- Reviewing work and improving SRE processes
- Collaborating with other teams outside Enterprise
Platforms
- Contributing to the strategic direction of the function
Skills And Qualifications
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related field.
- 5+ years of experience supporting production applications (ie SRE and/or DevOps roles)
- Practical understanding of implementing SLOs and SLIs
- Knowledge of Windows and/or Linux Systems administration and networking fundamentals
- Experience in implementing Observability and Alerting tools (eg Datadog, Splunk)
- Ability to automate application operations using tools such as Python, Java, Shell Scripting, Terraform, Chef, Puppet, SQL, Ansible
- Knowledge of AWS
- Experience in supporting middleware such as databases, webservers, MQ and Kafka
- Familiarity with containerization technologies, such as Docker and Kubernetes
- Excellent problem-solving skills and attention to detail
- Ability to work well under pressure and prioritize tasks in a fast-paced environment
- Fluency in English is essential.
- Ability to collaborate closely with others.
- Continual Improvement mindset
#LI-KA2
#LI-Hybrid
We are dedicated to fostering a collaborative, engaging, and inclusive environment and are committed to providing a workplace that empowers associates to be authentic and bring their best to work. We believe that associates do their best when they feel safe, understood, and valued, and we work diligently and collaboratively to ensure Broadridge is a company—and ultimately a community—that recognizes and celebrates everyone’s unique perspective.
Use of AI in Hiring
As part of the recruiting process, Broadridge may use technology, including artificial intelligence (AI)-based tools, to help review and evaluate applications. These tools are used only to support our recruiters and hiring managers, and all employment decisions include human review to ensure fairness, accuracy, and compliance with applicable laws. Please note that honesty and transparency are critical to our hiring process. Any attempt to falsify, misrepresent, or disguise information in an application, resume, assessment, or interview will result in disqualification from consideration.
Similar roles
-
Senior Site Reliability Engineer
GetFrankly Cluj-Napoca, Romania
-
SRE (Site Reliability Engineer) - H/F
Devoteam Levallois-Perret, Ile-de-France, France
-
Ingénieur Observabilité / SRE - H/F
Devoteam Levallois-Perret, Ile-de-France, France
-
SRE
Radware Tel-Aviv, Tel-Aviv District, Israel
-
SRE [Antifraud]
Plata Card Osnabrück, Lower Saxony, Germany
-
Senior SRE & Monitoring Developer
Ford Motor Company India