Senior Site Reliability Engineer (Copy) (Copy) (Copy)
Supio Seattle, Washington, United States · $200K–$240K/yr
Software Development · 51-200 employees
About the role
You will own the reliability and observability of the agentic platform while designing CI/CD pipelines to support rapid deployment. Additionally, you will partner with product engineering to strengthen core infrastructure and develop self-serve tooling.
What they look for
Requirements
Candidates must have 5+ years of experience in platform engineering, SRE, or infrastructure engineering with production-level ownership. Proficiency in AWS, Kubernetes, Terraform, and observability tooling is required, along with a bachelor's degree in a related field.
Benefits
Full description
Who we're looking for
We're looking for a senior platform engineer who thrives on solving problems that don't yet have a playbook. You'll work directly alongside Supio's core engineering team building our agentic product, helping shape how we develop, deploy, and observe AI agents in production — an area where the industry hasn't settled on standards yet. You're energized by ambiguity, comfortable moving between deep infrastructure work and close day-to-day collaboration with product engineers, and motivated by ownership rather than tickets. This is a hands-on, builder-oriented role for someone who wants to be embedded with the team shipping Supio's flagship agentic experience.
What you'll do
- Own Agent Reliability & Observability: Evolve our observability stack — metrics, logging, tracing, alerting — that keeps Supio's agentic platform reliable in production.Drive Agent CI/CD: Design and maintain the CI and release pipelines that let the core engineering team ship new agent capabilities quickly and safely, in a space where testing and deployment standards for AI agents are still being invented.
- Partner with Product Engineering: Work directly and daily alongside the core engineering team, embedding platform expertise into how they build.
- Strengthen Core Infrastructure: Support and improve the underlying infrastructure — AWS, Kubernetes, databases, networking — that the agentic platform runs on.
- Productize Self-Serve Tooling: Turn ad hoc infrastructure needs into self-serve tools the product team can use independently, reducing their dependency on direct infra support.
Qualifications
- Experience: 5+ years in platform engineering, SRE, or infrastructure engineering, with direct ownership of production systems.
- Education: Bachelor's degree in Computer Science or related field, or equivalent practical experience.
- Technical Skills: Production experience with AWS, Kubernetes, and infrastructure-as-code (Terraform); hands-on with CI/CD pipelines (GitHub Actions); working knowledge of databases (Postgres) and messaging/queuing systems ; strong networking fundamentals.
- Observability: Demonstrated experience building or integrating observability tooling (metrics, logging, tracing, alerting), including making build-vs-adopt decisions and driving self-serve adoption.
- AI/Agentic Platform Experience: Hands-on experience building, operating, or supporting an AI/LLM or agentic platform in production.
- Cross-functional Collaboration: Track record of working directly alongside product/application engineering teams.
- Automation & Scripting: Comfortable writing scripts or light code to automate infrastructure tasks, with a solid understanding of software architecture fundamentals.
- Judgment Under Ambiguity: Able to operate effectively without an established playbook and exercise sound technical judgment in a fast-changing environment.
Nice-to-haves
- Background in distributed systems, data pipelines, or data analysis
- Startup or early-stage company experience with a builder mindset
- Experience productizing internal tooling for self-serve use by other engineering teams
Salary
As an early-stage startup, we offer a competitive compensation package that includes base salary, meaningful equity, and benefits. Equity grants are designed to ensure employees share in the long-term success and upside of the company.
Base Salary by Location
As an early-stage startup, we offer a competitive compensation package that includes base salary, meaningful equity, and benefits. Equity grants are designed to ensure employees share in the long-term success and upside of the company.
Base Salary: $200,000 - $240,000 annually
Actual compensation may vary outside of these ranges based on a number of factors, including a candidate's qualifications, skills, competencies, experience, and geographic location.
Similar roles
-
Staff Site Reliability Engineer
Anduril Industries Costa Mesa, California, United States · $191K–$253K/yr
-
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Elastic Poland · PLN 359K–PLN 465K/yr
-
Junior Site Reliability Engineer
Semarchy United States
- Site Reliability Engineer
-
Analista Júnior em Infraestrutura | SRE
XP Inc. São Paulo, São Paulo, Brazil
-
Analista Pleno em Infraestrutura | SRE
XP Inc. São Paulo, São Paulo, Brazil