Senior Site Reliability Engineer (Cloud Platform)
Salve.Inno Consulting Philippines
Human Resources Services · 11-50 employees
About the role
The Senior Site Reliability Engineer will maintain the reliability, availability, and performance of production environments while driving automation and observability. They will also partner with software engineering teams to improve application reliability and participate in incident response and root cause analysis.
What they look for
Requirements
Candidates must have a Bachelor's or Master's degree in Engineering or Computer Science and strong experience with Kubernetes, AWS, and Linux environments. Proficiency in scripting languages like Python or Go and experience with Infrastructure as Code tools such as Terraform or Ansible are required.
Benefits
Full description
B2B Contract
Role Overview
We're looking for a Senior Site Reliability Engineer to help build, operate, and continuously improve a highly available cloud platform supporting mission-critical production services.
In this role, you'll work at the intersection of cloud infrastructure, software engineering, and operations, helping engineering teams build reliable, scalable systems while driving automation, observability, and operational excellence. You'll play a key role in strengthening platform reliability, improving incident response, and embedding SRE best practices throughout the software development lifecycle.
Key Responsibilities:
- Maintain the reliability, availability, and performance of production and pre-production environments.
- Monitor platform health and improve alerting, automation, and operational processes.
- Respond to production incidents, participate in root cause analysis, and implement long-term improvements.
- Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards.
- Partner with software engineers to improve application reliability throughout the development lifecycle.
- Develop and maintain operational documentation, troubleshooting guides, and runbooks.
- Automate repetitive operational tasks to improve efficiency and reduce manual intervention.
- Participate in on-call rotations while continuously improving incident response processes.
- Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams.
Requirements:
- Bachelor's or Master's degree in Engineering, Computer Science, or a related field.
- Strong experience operating Kubernetes or other container orchestration platforms.
- Experience supporting large-scale production services.
- Hands-on experience with AWS.
- Experience with Prometheus, Grafana, and ELK.
- Strong scripting skills (Bash, Python, or Go).
- Experience administering Linux-based production environments.
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible.
- Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing).
- Excellent troubleshooting, communication, and collaboration skills.
- A proactive mindset with a passion for automation and reliability.
Nice to Have:
- Experience with SIP or VoIP technologies.
- Familiarity with MySQL or PostgreSQL.
- Experience with Redis or other NoSQL databases.
What's on Offer:
- Long-term, full-time collaboration.
- Flexible remote working environment.
- Professional development opportunities, including training and technical learning.
- The opportunity to work on innovative cloud technologies used by customers worldwide.
- Collaborative engineering culture focused on knowledge sharing and continuous improvement.
- Modern Apple equipment provided.
Diversity and Inclusion Commitment
We are dedicated to creating and sustaining an inclusive, respectful workplace for all -regardless of gender, ethnicity, or background. We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.
Similar roles
-
Sovereign Engineering Platform SRE (m/f/d)
T-Systems Iberia Granada, Andalusia, Spain
-
Senior DevOps & Site Reliability Engineer (GCP)
Burjline Builders Sofia, Sofia-City, Bulgaria
- Platform Engineer (SRE - India)- III
-
Engineer Lead, Site Reliability
Zensar Pune, Maharashtra, India
-
Azure Site Reliability Engineer (SRE) - SaaS Operations
Zensar Bangalore South, Karnataka, India
-
Site Reliability Engineer
OSTTRA Gurugram, Haryana, India