Technical Support Lead (Senior Staff Engineer)
Crux AI Palo Alto, California, United States
About the role
You will build and lead the global technical support organization, establishing the company-wide customer support posture and case management architecture from scratch. You will serve as the primary technical bridge between enterprise customers and backend engineering teams to ensure reliable system performance.
What they look for
Requirements
The role requires 12+ years of experience in technical support or infrastructure operations, with at least 6 years in a leadership capacity. Candidates must possess deep expertise in large-scale compute environments and a proven track record of building support functions from the ground up.
Benefits
Full description
Built to set the gold standard for integrated AI infrastructure
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
Crux AI is led by CEO Ben Treynor Sloss, who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) discipline. At Crux AI, we apply Google’s SRE philosophy directly to customer operations: measuring reliability from the customer’s perspective, conducting blameless postmortems, and building systematic feedback loops that transform customer case escalations into permanent platform improvements.
WHAT YOU'LL DO
Crux AI is seeking a founding Technical Support Lead (Senior Staff Engineer) in Palo Alto, California to establish and own our company-wide customer technical support posture, case management architecture, and support technology stack from zero.
Reporting directly to the Chief Technical Officer (CTO), you will sit on the technical leadership team alongside the leads for Network Engineering, Hardware Operations, and Software/SRE. Serving as the primary connective tissue between our enterprise customers and backend engineering, you own the technical support doctrine, customer status communications during degradations, case escalation paths, and customer technical experience during critical issues.
We operate on a Direct-to-Expert support model. Rather than routing customers through generic support tiers, your team connects enterprise AI labs directly with senior technical support engineers who stay accountable for issue resolution end-to-end.
In this role, you will:
In this role, you will be responsible for building, managing, and leading the global technical support organization. However, this is a true hands-on, builder seat, not a supervisory position early on. In the initial phase, you will handle the first complex case escalations, triage cluster issues alongside SRE, and build the support runbooks and status workflows yourself before building out your team of engineers and managers. You must revel in ambiguity, act with high autonomy, and maintain a strict "discipline over heroics" mindset, building automated support systems and 24x7 coverage that make handling the next customer escalation routine.
Key Responsibilities:
- Define Direct-to-Expert Case Architecture: Design, roll out, and govern the company-wide case triage framework, Direct-to-Expert ownership model led by senior technical support engineers, escalation boundaries, and severity definitions across Support, SRE, Network Engineering, and Google Support.
- Proactive Account Technical Management (TAM) & Onboarding: Provide proactive technical guidance for strategic enterprise accounts during initial cluster bring-up, model training spin-ups, job placement, and checkpointing strategies to prevent issues before they occur.
- Connective Tissue to Product & Engineering: Systematically aggregate ticket analytics, case trends, and postmortem findings to feed prioritized product bug fixes and platform feature requests directly to SRE, Network Engineering, and Product teams.
- Own Customer Status Communications & Support Stack: Own company status page messaging, incident notification channels, and customer-facing status updates during cluster degradations, while governing the selection and deployment of the support ticketing technology stack.
- Build & Scale 24x7 Support Operations: Recruit, structure, and manage a high-bar technical support engineering organization capable of delivering robust 24x7 operational coverage, establishing shift leads and escalation tiers as fleet capacity ramps.
- Embed AI-Agentic First Support: Deploy AI agents and LLM-driven diagnostic workflows as first principles across case triage, log parsing, root-cause analysis, and customer updates before scaling headcount.
- Customer SLA Alignment & Executive Reporting: Partner with Sales, Legal, SRE, and Product to ensure support workflows strictly align with contractual customer SLAs, delivering monthly executive reporting on case resolution velocity, friction trends, and customer health.
SIGNALS OF SUCCESS
After 60 days in this role:
- Mastered Crux AI's technical stack and internal escalation paths; published the v1 Direct-to-Expert case triage framework and severity model to SRE, Network Engineering, and Hardware Operations; launched initial customer status page & messaging protocols.
After 6 months:
- Customer case escalation framework active company-wide; automated feedback loop delivering case analytics to Product and SRE; 24x7 shift coverage established; initial senior technical support engineers and shift managers onboarded.
After 1 year:
- Support organization operating against published customer SLAs with board-visible customer health dashboards; formal TAM program running for top accounts; >90% of routine case triage workflows handled via automated agentic pipelines.
EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH
Experiences
- 12+ Years Experience: 12+ years in technical support, TAC, or customer-facing infrastructure/SRE operations for large-scale compute/cloud environments, including 6+ years managing technical support engineering teams and leaders.
- Deep Hyperscale Triage Depth: Hands-on ability to triage and resolve complex customer issues across high complexity large scale environments by analyzing logs, alerts, and telemetry data, through grounding in an architectural understanding of complex systems that extends far beyond executing static playbooks.
- Direct-to-Expert Leadership Experience: Track record building or operating high-touch support organizations for complex enterprise infrastructure, moving away from generic multi-tier queues toward direct-ownership models anchored by senior technical support engineers.
- Senior Technical Leadership Scope: Proven track record building a technical support function from scratch, operating as a peer to senior engineering directors, and presenting customer health metrics to executive teams.
Attributes
- Customer-Centric Outcome Orientation: Treats the customer’s workload execution as the sole definition of success, recognizing that a case marked "resolved" while a customer's multi-thousand-node training job is still stalled is not resolved.
- Discipline Over Heroics: Rejects "heroic all-nighters" in favor of systematic postmortems, automated runbooks, and robust escalation architectures.
Mindset
- Builder Mindset & Ambiguity: High tolerance for ambiguity; comfortable operating as a hands-on lead early on, establishing order amidst rapid fleet growth, and designing scalable 24x7 organizational systems.
- AI-Agentic First: Active utilization of AI agents and automated diagnostic tools in technical operations, defaulting to agentic workflows over manual process.
Nice to have (Preferred, not required):
- Hyperscaler / Neocloud Leadership: Technical support or TAC leadership at a hyperscaler (Google, AWS, Meta, Microsoft) or leading neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid customer growth.
- Direct TPU / Frontier Silicon Support: Direct experience supporting Google TPU clusters or large-scale GPU accelerator pods in production.
- Contractual SLA Mechanics: Familiarity with customer SLA terms, support tiering, maintenance windows, and service credit mechanics in enterprise cloud contracts.
- Support Tooling & Budget Ownership: Financial accountability for support platforms (Zendesk/ServiceNow/custom), status page tooling, vendor contracts, and global 24x7 shift coverage.
Salary Range Information
The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Crux
- We offer generous base, bonus and additional incentive based compensation
- Health, dental, and vision coverage for you and your dependents
- Company-paid life insurance and disability
- Full suite of other optional benefits
- 401(k) Plan with 4% company match
- Hybrid schedule offering four days in-office collaboration paired with one remote workday for focused, individual work