JuiceX

Senior Site Reliability Engineer

JuiceX Australia

Software Development · 11-50 employees

8 h ago
sre Senior (5-10 yrs) Contractor Ireland
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the reliability, scalability, and operational integrity of the Exchange's cloud and Kubernetes infrastructure. Responsibilities include managing infrastructure as code, leading incident response, and implementing policy as code to ensure regulatory compliance.

What they look for

GCP Kubernetes Terraform GitOps Flux Argo Observability Prometheus OpenTelemetry Rust AlloyDB Postgres ClickHouse Redis MongoDB Incident Command

Requirements

Candidates must have 5+ years of experience in SRE or infrastructure engineering within a regulated environment. Proficiency in GCP, Kubernetes, Terraform, and GitOps is required, along with a willingness to participate in a 24/7 on-call rotation.

Full description

The Opportunity

We are hiring Senior Site Reliability Engineers to own the reliability, scalability, and operational integrity of JuiceX's Exchange infrastructure. As a CFTC regulated market, our platform must run continuously and meet exacting standards for uptime, performance, and audibility. You will work alongside Engineering, Security, and Compliance to keep the Exchange running in a high stakes, regulated environment.

Professional Experience

  • Production experience with GCP (or AWS with a credible path to GCP) and Kubernetes
  • Terraform at estate scale: module design, state management, migrations, drift, and blast radius
  • GitOps with Flux or Argo, including progressive delivery
  • Observability design: SLOs, error budgets, distributed tracing, and Prometheus or OpenTelemetry
  • Incident command during production outages
  • Reading Rust services and reasoning about their failure modes
  • Database operations at scale: AlloyDB or Postgres, ClickHouse, Redis, MongoDB

Qualifications

  • 5+ years in SRE, DevOps, or production infrastructure engineering
  • Bachelor's degree in a technical field, or equivalent experience
  • Experience in financial services, trading, clearing, or another regulated environment where the audit trail is part of the product
  • Willingness to work a 24/7 on call rotation, including nights, weekends, and holidays
  • Strong written and verbal communication across Engineering, Security, and Compliance
  • Preferred: policy as code (OPA/Rego, Kyverno, Checkov) and supply chain integrity (SHA and digest pinned images, provenance)
  • Preferred: cross cloud networking (VPN, PrivateLink, split horizon DNS, service mesh)
  • Preferred: load and capacity testing (k6 or equivalent)

Position Responsibilities

  • Own reliability, availability, and performance of the Exchange's cloud and Kubernetes infrastructure
  • Maintain infrastructure as code in Terraform, managing state, drift, and blast radius
  • Operate GitOps deployment pipelines with progressive delivery and safe rollback
  • Define and maintain SLOs, error budgets, and observability
  • Lead incident response and post incident reviews as incident commander during outages
  • Build automation and guardrails for safe deployment in a regulated environment
  • Implement policy as code and supply chain integrity controls for regulatory and audit compliance
  • Design and regularly test multi region and disaster recovery capabilities
  • Partner with Engineering, Security, and Compliance on platform security and resilience

Similar roles