Carbon3.ai

Site Reliability Engineer

Carbon3.ai United Kingdom

Technology, Information and Internet · 11-50 employees

Jun 10
sre Mid (2-5 yrs) Full-time United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build Python-based automation for incident triage, runbook execution, and routine operational tasks to improve reliability. Additionally, you will integrate observability and infrastructure APIs to enhance alert quality and develop internal tools for self-service capabilities.

What they look for

Python Site Reliability Engineering Platform Engineering Automation Observability Prometheus Grafana Incident Management Runbook Automation API Integration ChatOps Infrastructure as Code Monitoring ITSM Distributed Tracing OpenTelemetry

Requirements

The role requires experience in SRE, Platform Engineering, or production infrastructure operations, along with hands-on experience in observability tooling. Candidates must also possess proficiency in Python for automation and have experience converting manual runbooks into automated workflows.

Full description

Role Summary:

We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response.

Key Responsibilities:

  • Build Python-based automation for incident triage, runbook execution, and routine operational tasks.
  • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
  • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
  • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
  • Maintain version-controlled runbook-as-code and automation libraries.
  • Translate post-incident learnings into better tooling, automation and operational standards.
  • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

Essential Experience:

  • Experience in SRE, Platform Engineering, or production infrastructure operations.
  • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
  • Exposure to incident management / on-call and converting manual runbooks into automation.
  • Experience with Python for automation, APIs and integrations.

Nice To Have:

  • GPU, datacentre or colocation infrastructure experience.
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
  • ChatOps tooling (Slack or Microsoft Teams bots).
  • OpenTelemetry, logging or distributed tracing experience.
  • DCIM, IPAM or hypervisor-control-plane integrations.
  • Experience with LLM-assisted or agent-based operational automation.

Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

Diversity & Inclusion:

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar roles