About the role
You will own the reliability, scalability, and operational integrity of the Exchange's cloud and Kubernetes infrastructure. Responsibilities include managing infrastructure as code, leading incident response, and implementing policy as code to ensure regulatory compliance.
What they look for
Requirements
Candidates must have 5+ years of experience in SRE or infrastructure engineering within a regulated environment. Proficiency in GCP, Kubernetes, Terraform, and GitOps is required, along with a willingness to participate in a 24/7 on-call rotation.
Full description
The Opportunity
We are hiring Senior Site Reliability Engineers to own the reliability, scalability, and operational integrity of JuiceX's Exchange infrastructure. As a CFTC regulated market, our platform must run continuously and meet exacting standards for uptime, performance, and audibility. You will work alongside Engineering, Security, and Compliance to keep the Exchange running in a high stakes, regulated environment.
Professional Experience
- Production experience with GCP (or AWS with a credible path to GCP) and Kubernetes
- Terraform at estate scale: module design, state management, migrations, drift, and blast radius
- GitOps with Flux or Argo, including progressive delivery
- Observability design: SLOs, error budgets, distributed tracing, and Prometheus or OpenTelemetry
- Incident command during production outages
- Reading Rust services and reasoning about their failure modes
- Database operations at scale: AlloyDB or Postgres, ClickHouse, Redis, MongoDB
Qualifications
- 5+ years in SRE, DevOps, or production infrastructure engineering
- Bachelor's degree in a technical field, or equivalent experience
- Experience in financial services, trading, clearing, or another regulated environment where the audit trail is part of the product
- Willingness to work a 24/7 on call rotation, including nights, weekends, and holidays
- Strong written and verbal communication across Engineering, Security, and Compliance
- Preferred: policy as code (OPA/Rego, Kyverno, Checkov) and supply chain integrity (SHA and digest pinned images, provenance)
- Preferred: cross cloud networking (VPN, PrivateLink, split horizon DNS, service mesh)
- Preferred: load and capacity testing (k6 or equivalent)
Position Responsibilities
- Own reliability, availability, and performance of the Exchange's cloud and Kubernetes infrastructure
- Maintain infrastructure as code in Terraform, managing state, drift, and blast radius
- Operate GitOps deployment pipelines with progressive delivery and safe rollback
- Define and maintain SLOs, error budgets, and observability
- Lead incident response and post incident reviews as incident commander during outages
- Build automation and guardrails for safe deployment in a regulated environment
- Implement policy as code and supply chain integrity controls for regulatory and audit compliance
- Design and regularly test multi region and disaster recovery capabilities
- Partner with Engineering, Security, and Compliance on platform security and resilience
Similar roles
-
DevOps & Site Reliability Engineer (GCP)
Burjline Builders Tbilisi, Georgia
-
Sr. Site Reliability Engineer - Top Secret Clearance (Starlink)
SpaceX Redmond, Washington, United States · $165K–$230K/yr
-
Senior Staff Site Reliability Engineer
Ping Identity Denver, Colorado, United States · $170K–$227K/yr
-
JD IC4 - Sr Infra Engineer - SRE
Spin Careers Hamburg, Germany
-
Senior Site Reliability Engineer
ADT Whitpain Township, Pennsylvania, United States · $129K–$193K/yr
-
[8SN] Site Reliability Engineer (SRE) – UI/UX
Software Mind Montreal, Quebec, Canada