PCI Pal

Site Reliability Engineer

PCI Pal United Kingdom

Information Technology & Services · 51-200 employees

4 d ago
sre Senior (5-10 yrs) Full-time United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will act as the technical authority for observability and reliability, managing DataDog implementations and infrastructure monitoring. You will also lead incident response, automate manual toil, and collaborate with development teams to improve system resiliency and capacity planning.

What they look for

DataDog Observability Reliability Engineering Infrastructure Monitoring Capacity Planning Incident Response Automation SLIs SLOs Networking Telephony Canary Deployments Jenkins Jira PCI Compliance Disaster Recovery

Requirements

The role requires extensive hands-on experience with observability tools like DataDog and a strong background in incident response and infrastructure monitoring. Candidates should possess solid software engineering skills and experience with compliance frameworks, capacity planning, and automated build systems.

Benefits

25 days holiday Medical insurance Dental insurance Optical insurance Flexible working environment Electric Vehicle Scheme Work from anywhere policy Reward and benefits hub Training and development opportunities

Full description

You'll join PCI Pal's DevOps team as the go-to authority on observability and reliability, owning our DataDog implementation and extending monitoring, alerting and anomaly detection across infrastructure, networking and telephony. Working closely with the senior DevOps engineer and the wider team, you'll drive down manual toil, improve resiliency, and own capacity planning, availability targets and disaster recovery readiness as load grows. You're joining an existing team and will be relied on to raise the reliability and observability bar for the whole group.

  • Extensive, hands-on experience with DataDog (or equivalent tools such as New Relic, Grafana/Prometheus, Dynatrace)
  • Experience managing and controlling observability tooling costs at scale
  • Solid experience monitoring networking infrastructure (latency, packet loss, device health, anomaly detection)
  • Experience with capacity planning, demand forecasting, and redundancy/failover design
  • Strong background in incident response practices and building automation around them
  • Experience defining SLIs/SLOs and error-budget-based alerting
  • Solid software engineering skills, writing maintainable, testable, version-controlled automation and tooling
  • Experience implementing progressive delivery practices (canary deployments, staged rollouts, automated rollback)
  • Confidence engaging with and influencing development teams
  • Experience of PCI compliance or other similar compliance frameworks
  • Experience with automated build systems (e.g. Jenkins), work management systems (e.g. Jira), and source control

Nice to haves:

  • Experience with telephony/voice infrastructure monitoring
  • Experience mentoring or upskilling other engineers
  • Own the observability framework and standards across infrastructure, networking and telephony, defining what good monitoring and alerting look like, and partnering with development teams to implement against these standards, including a practical monitoring scorecard, reviewing and signing off on their coverage
  • Control observability costs - right-sizing usage (hosts, custom metrics, log ingestion/retention, APM) and making deliberate cost-vs-coverage tradeoffs
  • Work with teams to reduce alert noise while making sure genuine incidents are still surfaced immediately
  • Act as the technical authority on monitoring maturity across the engineering function
  • Mentor and upskill the wider DevOps team on reliability and observability practices, working alongside the senior DevOps engineer
  • Identify and automate away manual toil, including incident response workflows, compliance evidence collection and release processes
  • Drive resiliency improvements across the infrastructure estate - redundancy, failover and capacity planning
  • Lead on incident response and resolution, contributing to postmortem practices that turn failures into lasting fixes
  • Work with development teams to embed resiliency practices into release processes, e.g. progressive rollouts, canary deployments and fast rollback, so changes can ship quickly without risking availability
  • 25 days holiday, rising to 28 days per annum with length of service
  • Medical, dental and optical insurance cover
  • Option to either work in our Ipswich office, or from home (or both!)
  • An exciting and flexible working environment surrounded by friendly and committed co-workers
  • Electric Vehicle Scheme incentive
  • “Work from anywhere” 2 weeks per year policy
  • Reward, benefits and wellbeing hub (offering support, discounts, cashback and savings)
  • Training and development opportunities
  • Ad-hoc team events, incentives and competitions

Similar roles