Senior Site Reliability Engineer
Kontakt.io New York, New York, United States · $200K–$240K/yr
Technology, Information and Internet · 51-200 employees
About the role
You will design, build, and operate resilient, self-healing infrastructure on AWS while managing incident response and disaster recovery protocols. Additionally, you will maintain CI/CD pipelines and ensure the platform meets strict security and compliance standards for healthcare data.
What they look for
Requirements
Candidates must have at least 4 years of experience in Site Reliability Engineering or Cloud Infrastructure with deep expertise in AWS and Kubernetes. A proven track record in incident management, infrastructure automation, and working within regulated environments is essential.
Benefits
Full description
About Kontakt.io
Inside health systems, where every second can matter, operations are still spread across dozens of disconnected tools and platforms. Kontakt.io is changing that.
We combine proprietary hardware, AI-powered intelligence, and deep integrations with the technology health systems already have in place to build real-time understanding of what's happening across their operations. That intelligence becomes the execution layer care teams have been missing, helping them make smarter decisions and deliver better patient care.
Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we’ve more than doubled our revenue and are rapidly scaling with a clear path toward $100M in annual recurring revenue.
If you're excited to solve hard problems and help health systems deliver better care, we'd love to meet you!
About the role
We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their hands directly into the systems that keep our healthcare platform running for hospitals and care teams who can't afford downtime. This is a builder's seat — you'll carry real operational weight and have direct influence over how our infrastructure evolves.
What you'll do
- Personally design, build, and operate resilient, self-healing infrastructure across our AWS-based platform
- Own incident response end-to-end: detection, mitigation, root-cause investigation, and postmortems that actually change how the system behaves next time
- Design and run disaster-recovery and failover exercises with real RTO/RPO targets — you'll be the one who knows exactly what happens when things break
- Build out observability that's genuinely tuned — SLIs, SLOs, and alerting people trust, not noise
- Build and maintain CI/CD pipelines and infrastructure as code (Terraform, GitOps)
- Work hands-on in Kubernetes, below the abstraction layer — you'll know the system, not just the dashboard
- Partner day-to-day with our platform lead, sharing real production ownership and on-call
- Shape our security and compliance posture (HIPAA, SOC 2 Type 2) as it relates to infrastructure handling protected health data
What you bring
- 4+ years in Site Reliability Engineering or Cloud Infrastructure
- Deep, current expertise in AWS, Kubernetes, and distributed systems, with the depth to go past the vocabulary
- Real experience running disaster recovery or failover exercises
- A track record of driving incident response and postmortems yourself
- Solid grounding in CI/CD automation, GitOps, and infrastructure as code
- An appetite for staying close to the system rather than one step removed from it
- Bonus: healthcare IT, EHR data, or HIPAA/SOC 2-governed environments
- Bonus: experience with high-traffic, mission-critical SaaS or IoT platforms
Logistics, Perks & Benefits
- Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office
- Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors
- Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work
- Autonomy to solve meaningful problems with work that ships quickly and makes a difference
Compensation
The expected salary range for this role is $200,000 – $240,000 for New York-based candidates. Actual compensation within this range will be determined based on relevant experience, skills, and qualifications. In exceptional cases, where a candidate’s experience or qualifications significantly exceed those anticipated for this role, we may consider the candidate for a more senior level. This role may also be eligible for equity and bonus compensation.
Similar roles
-
Senior Site Reliability Engineer
MetaRouter $160K–$180K/yr
-
Senior Application Support Engineer / Site Reliability Engineer (SRE)
DTCC Boston, Massachusetts, United States
-
Site Reliability Engineer - Privilege & Access Management
JPMorgan Chase & Co. Plano, Texas, United States
-
Senior Site Reliability Developer
Vena Solutions Canada · CA$123K–CA$167K/yr
-
Manager TAG and Encompass SRE
WEXWEXUS Berwyn, Illinois, United States · $106K–$131K/yr
-
CTIO - Site Reliability Engineer- Senior Associate
PwC Birmingham, Alabama, United States · $55K–$187K/yr