Site Reliability Engineer
PCI Pal United Kingdom
Information Technology & Services · 51-200 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
You will act as the technical authority for observability and reliability, managing DataDog implementations and infrastructure monitoring. You will also lead incident response, automate manual toil, and collaborate with development teams to improve system resiliency and capacity planning.
What they look for
Requirements
The role requires extensive hands-on experience with observability tools like DataDog and a strong background in incident response and infrastructure monitoring. Candidates should possess solid software engineering skills and experience with compliance frameworks, capacity planning, and automated build systems.
Benefits
Full description
You'll join PCI Pal's DevOps team as the go-to authority on observability and reliability, owning our DataDog implementation and extending monitoring, alerting and anomaly detection across infrastructure, networking and telephony. Working closely with the senior DevOps engineer and the wider team, you'll drive down manual toil, improve resiliency, and own capacity planning, availability targets and disaster recovery readiness as load grows. You're joining an existing team and will be relied on to raise the reliability and observability bar for the whole group.
- Extensive, hands-on experience with DataDog (or equivalent tools such as New Relic, Grafana/Prometheus, Dynatrace)
- Experience managing and controlling observability tooling costs at scale
- Solid experience monitoring networking infrastructure (latency, packet loss, device health, anomaly detection)
- Experience with capacity planning, demand forecasting, and redundancy/failover design
- Strong background in incident response practices and building automation around them
- Experience defining SLIs/SLOs and error-budget-based alerting
- Solid software engineering skills, writing maintainable, testable, version-controlled automation and tooling
- Experience implementing progressive delivery practices (canary deployments, staged rollouts, automated rollback)
- Confidence engaging with and influencing development teams
- Experience of PCI compliance or other similar compliance frameworks
- Experience with automated build systems (e.g. Jenkins), work management systems (e.g. Jira), and source control
Nice to haves:
- Experience with telephony/voice infrastructure monitoring
- Experience mentoring or upskilling other engineers
- Own the observability framework and standards across infrastructure, networking and telephony, defining what good monitoring and alerting look like, and partnering with development teams to implement against these standards, including a practical monitoring scorecard, reviewing and signing off on their coverage
- Control observability costs - right-sizing usage (hosts, custom metrics, log ingestion/retention, APM) and making deliberate cost-vs-coverage tradeoffs
- Work with teams to reduce alert noise while making sure genuine incidents are still surfaced immediately
- Act as the technical authority on monitoring maturity across the engineering function
- Mentor and upskill the wider DevOps team on reliability and observability practices, working alongside the senior DevOps engineer
- Identify and automate away manual toil, including incident response workflows, compliance evidence collection and release processes
- Drive resiliency improvements across the infrastructure estate - redundancy, failover and capacity planning
- Lead on incident response and resolution, contributing to postmortem practices that turn failures into lasting fixes
- Work with development teams to embed resiliency practices into release processes, e.g. progressive rollouts, canary deployments and fast rollback, so changes can ship quickly without risking availability
- 25 days holiday, rising to 28 days per annum with length of service
- Medical, dental and optical insurance cover
- Option to either work in our Ipswich office, or from home (or both!)
- An exciting and flexible working environment surrounded by friendly and committed co-workers
- Electric Vehicle Scheme incentive
- “Work from anywhere” 2 weeks per year policy
- Reward, benefits and wellbeing hub (offering support, discounts, cashback and savings)
- Training and development opportunities
- Ad-hoc team events, incentives and competitions
Similar roles
-
Senior Site Reliability Engineer (SRE)
LeoLabs, Inc. $171K–$192K/yr
-
Senior Site Reliability Engineer
2K Austin, Texas, United States
-
Lead Site Reliability Engineer
Sherwin-Williams Cleveland, Ohio, United States
-
Senior Site Reliability Engineer (SRE)
Tradeweb United States · $170K–$210K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google New York, New York, United States · $174K–$252K/yr
-
Software Developer III, Site Reliability
Google Waterloo, Ontario, Canada · CA$150K–CA$153K/yr