Senior Site Reliability & Software Engineering Manager
Crux AI Palo Alto, California, United States
About the role
You will build and lead a team of senior engineers to manage fleet reliability, focusing on automating remediation and eliminating manual pages. You will also own the software lifecycle for bare-metal provisioning, control plane reliability, and the integration of AI/ML techniques into operational workflows.
What they look for
Requirements
Candidates must have 10+ years of software or infrastructure engineering experience with at least 3 years in a management role. A strong technical background in SRE principles, distributed systems, and hands-on coding in Go, Python, or C++ is required.
Benefits
Full description
Built to set the gold standard for integrated AI infrastructure
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
Crux AI is led by CEO Ben Treynor Sloss, who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) discipline. At Crux AI, we treat operations fundamentally as a software engineering problem.
WHAT YOU'LL DO
We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA.
In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.
This is a true hands-on, builder seat, not a supervisory position. In the early days, you will write the first remediation controllers, set reliability baselines, and take initial pages yourself to stay close to the work before expanding your team. You will lead an elite group of unusually senior software and reliability engineers: engineers with significantly greater technical and software depth than traditional operational SRE orgs. You must thrive on independence and revel in ambiguity, turning unknowns into concrete engineering priorities in a fast-paced, high-growth environment. While you will devise and participate in initial on-call rotations, your core mandate is to combine SRE disciplines with extensive AI/ML automation to drive operational pages down to zero.
In this role, you will:
- Build & Lead Senior Engineering Teams: Recruit, lead, and mentor an initial team of senior software and reliability engineers across Palo Alto and Europe as fleet capacity ramps rapidly.
- Automate Pages to Zero: Own fleet availability end-to-end; establish on-call rotations while relentlessly developing self-healing systems and predictive remediation to eliminate manual pages.
- Embed AI/ML into SRE Disciplines: Apply agentic techniques, machine learning models, and automated diagnostic workflows extensively to telemetry collection, root-cause analysis, and predictive cluster recovery.
- Own Bare-Metal & Fleet Lifecycle Software: Drive software engineering for bare-metal node provisioning, firmware deployment, thermal/stress burn-in validation, host/TPU health monitoring, and decommissioning.
- Control Plane & Fabric Reliability: Own software reliability for cluster orchestration, scheduling, capacity allocation APIs, and high-performance TPU host/interconnect networks.
- Define Observability & SLOs: Establish customer-facing SLIs/SLOs (job goodput, time-to-detect, node availability) and build the telemetry pipelines serving operators, executives, and customers.
SIGNALS OF SUCCESS
After 60 days in this role:
- Baseline reliability framework defined (v1 SLOs, severity structure, change management); first AI/ML-driven automated remediation controller shipped to production; recruiting active for senior engineering hires in Palo Alto.
After 6 months:
- First TPU cluster brought online under your team’s automated acceptance criteria; automated remediation pipeline running with pass rates tracked; observability v1 in daily production use; core senior team onboarded
After 1 year:
- TPU fleet operating against published customer SLOs; >90% of node/fabric faults automatically quarantined and remediated without human paging; team scaled ahead of rapid capacity ramps.
EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH
Experiences
- 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.
- Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset.
- Hands-On Technical Depth (SRE + SWE): Track record shipping production code in Go, Python, or C++, with hands-on systems expertise across Linux OS kernels, bare-metal provisioning, firmware, and/or high-performance networking fabrics. Extensive experience with distributed systems and open source software.
- Builder Mindset & Ambiguity: A true "builder, not supervisory" orientation; comfortable operating with high autonomy, navigating ambiguity, and establishing structure amidst rapid growth.
- Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.
- Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments.
Attributes
- Possess a high tolerance for ambiguity. The first clusters will carry customer workloads while the SLOs are still being defined and the team is still being hired. You absorb that, translate unknowns into concrete near-term priorities, and never manufacture false certainty about reliability the data does not support.
- Understands that the customer’s job is the unit of reliability. A node that is “up” while a training run stalls on a flapping link is down. You measure what customers experience — goodput, time-to-recover, lost progress — and hold the whole stack, and Google, to it.
Mindset
- AI-agentic first. Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.
Nice to have (Preferred, not required):
- Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.
- Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).
- Custom Fleet Tooling: Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.
- Facility Telemetry & Thermal Signals: Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.
- Customer SLAs & Reporting: Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews.
Salary Range Information
The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Crux
- We offer generous base, bonus and additional incentive based compensation
- Health, dental, and vision coverage for you and your dependents
- Company-paid life insurance and disability
- Full suite of other optional benefits
- 401(k) Plan with 4% company match (USA employees)
Similar roles
-
Staff Site Reliability Engineer
KEV Group Toronto, Ontario, Canada · $150K–$180K/yr
-
Site Reliability Engineer - Data Platform
IMC Amsterdam, North Holland, Netherlands
-
Staff Software Engineer, Site Reliability Engineering, Vertex AI
Google Warsaw, Masovian Voivodeship, Poland · PLN 480K–PLN 492K/yr
-
Site Reliability Engineering Manager
Conifers.ai Tel-Aviv, Tel-Aviv District, Israel
-
Senior Site Reliability Engineer
Precisely International Jobs Bielsko-Biała, Silesian Voivodeship, Poland
-
Network SRE
JPMorgan Chase & Co. Buenos Aires, Argentina