About the role
You will own the production infrastructure, including deploys, reliability, observability, and cost management for agentic AI systems. You are responsible for building the tooling to enforce reliability and maintaining the security posture required for enterprise-grade compliance.
What they look for
Requirements
You must have prior experience owning production infrastructure at a small company and possess deep expertise in AWS, containers, and infrastructure as code. You should be pragmatic about scaling and have a strong background in observability and enterprise security standards.
Full description
The opportunity
We're building at the frontier of what applied AI can do inside real work — the operational core of an industry that moves the physical economy. One of the largest on earth, essential, and almost entirely untouched by modern AI. It still runs on people, spreadsheets, and software written before the internet.
We went vertical first on purpose. Vertical is where the hard problems live: messy inputs, real consequences for being wrong, decades of institutional knowledge nobody wrote down. Anything that works here has been tested against reality in a way horizontal tooling never is.
And it doesn't stay vertical. The systems we're building — how work gets decomposed, verified, corrected, and learned from — aren't specific to one industry. They're specific to work. The vertical is the proving ground. The reapplication is the company.
We're deliberately quiet about which industry until we talk. What we'll say now: it's enormous, it's overlooked, and the incumbents aren't coming. You'll get the full picture on the first call.
Why this might be interesting
- A cap table most early-stage companies would envy. Raised at the top quartile of seed-stage rounds by size — 36+ months of runway — from the seed investors who were early in Palantir, Databricks, Anduril, GitLab, Retool, Lyft, Square, DoorDash, Superhuman, and Ironclad.
- Real traction, right now. Live in production with design partners, and the data flywheel is already turning. Demand isn't the bottleneck — execution is. The market is tens of thousands of enterprises, each worth seven figures a year.
- Founded by multi-time exited operators. The founding team has built and sold multiple software companies and spent years up close with dozens more. You're joining people who know how this is actually done — not learning it alongside them.
The seat
You own how we run. Deploys, environments, reliability, observability, cost, and the security posture that lets enterprises say yes.
Here's why this isn't a standard platform seat. The workloads are strange. Agentic systems are long-running where web services are fast, bursty where traffic is smooth, expensive per unit of work in a way that makes cost an architectural concern rather than a finance one, and non-deterministic in a way that breaks most of what "healthy" means on a dashboard. A green status page can sit on top of a system quietly producing wrong answers. Standard observability tells you the call succeeded. It won't tell you the output was garbage.
Figuring out what reliability even means for this class of system — then building the tooling to enforce it — is genuinely unsolved. That's the seat.
- Infrastructure as code, environments, and deploys that a small team can move fast on without breaking production.
- Observability built for probabilistic systems — tracing across async and long-running work, and the harder problem of surfacing quality regressions, not just failures.
- Cost and performance as first-class engineering constraints.
- The security and compliance foundation enterprise buyers diligence before they sign.
Who you are
- You've owned production infrastructure at a small company. (Hard requirement.) You were the person who got paged, and the person who fixed the thing that caused the page — not one specialist layer inside a large platform org.
- Deep on AWS and infrastructure as code. Containers, CI/CD, IaC — plus the judgment to know what a seed-stage company should not build yet.
- Observability is a craft to you, not a vendor. You've instrumented systems where the interesting failures were the quiet ones.
- You've been close to enterprise security review. SOC 2, access control, secrets, data handling. You know what buyers ask and how to be ready before they ask.
- Pragmatic about scale. You right-size for today with an honest read on what breaks at 10x, and you know the difference between necessary foundations and premature infrastructure.
The stack
AWS · Python · Docker · Terraform · PostgreSQL · Redis · Celery
Referral bounty
Know someone exceptional? Introduce us. If we hire them into a full-time role, we'll send you $5,000.
Similar roles
-
Senior DevOps Engineer
realworld-one Bengaluru, Karnataka, India
-
Senior Data Platform DevOps Engineer (m/w/d) - Azure & Lakehouse (Ref.Nr.: 47816)
Wavestone Germany AG Zurich, Zurich, Switzerland
-
DevOps Engineer
Unicon Systems Hillsborough County, Florida, United States · $125K–$135K/yr
-
Lead DevOps Engineer
myenergi United Kingdom
-
DevOps Engineer
Eurofins Belo Horizonte, Minas Gerais, Brazil
-
DevOps Engineer-II
CommerceIQ Bengaluru, Karnataka, India