TailorCare

Senior Site Reliability Engineer

TailorCare United States

Hospitals and Health Care · 51-200 employees

6 h ago
Remote sre Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and scale infrastructure by implementing standardized AWS footprints, universal CI/CD pipelines, and observability tools. Additionally, you will act as an engineering advocate to support product teams, manage on-call rotations, and ensure compliance with security standards.

What they look for

Site Reliability Engineering AWS Terraform CI/CD Python Go TypeScript Observability DORA metrics Infrastructure as Code HIPAA compliance HITRUST compliance Incident management Automation System architecture

Requirements

Candidates must have 5+ years of experience in SRE, DevOps, or Software Engineering with deep expertise in AWS and Terraform. Strong programming skills and experience with modern observability stacks and CI/CD methodologies are required.

Benefits

Medical insurance Dental insurance Vision insurance Life insurance Disability insurance Wellness resources Employer HSA contribution 401k plan with employer matching Paid time off Paid parental leave

Full description

About TailorCare

TailorCare is transforming the experience of specialty care. Our comprehensive care program takes a profoundly personal, evidence-based approach to improving patient outcomes for joint, back, and muscle conditions. By carefully assessing patients' symptoms, health histories, preferences, and goals with predictive data and the latest evidence-based guidelines, we help patients choose and navigate the most effective treatment pathway for them every step of the way.

TailorCare values the experiences and perspectives of individuals from all backgrounds. We are a highly collaborative, curious, and determined team passionate about scaling a high-growth start-up to improve the lives of those in pain. TailorCare is a remote-first company with our corporate office located in Nashville.

About the Role

TailorCare is building a world-class engineering culture, and we are looking for a Senior Site Reliability Engineer to help build and scale our infrastructure. As an early hire in our newly formed Infrastructure & SRE department, you will work directly with the Director of Infrastructure & SRE, who sets our technical direction, and you will play a hands-on role in translating that vision into a resilient, scalable, and secure platform. You will have opportunities to shape how the work gets done and to influence technical decisions through strong execution and informed recommendations.

This is not a traditional "ops" role where you just close tickets. We are looking for an empathetic engineering advocate who partners closely with product engineering teams. You will help us execute on major initiatives in 2026 and 2027, including rolling out universal CI/CD systems, standardizing Terraform, tracking and driving improvement of DORA metrics, leveraging AI for automation, and establishing SLOs/SLIs and error budgets. If you love building paved roads for developers and bringing calm, automated discipline to production environments, this is the role for you.

Primary Responsibilities

  • Implement the standardization of our AWS footprint using Terraform. Eliminate manual provisioning, establish reproducible environments, and integrate AI-assisted tooling (e.g., automated PR reviews, intelligent infrastructure operations) to accelerate our workflows.
  • Build and maintain universal CI/CD pipelines (e.g., GitHub Actions) that remove friction from the SDLC. Track and improve DORA metrics across engineering teams to ensure we are shipping quickly and safely.
  • Implement observability improvements across AWS and third-party integrations. Monitor and support SLOs, SLIs, and error budgets for key services, ensuring high availability for our web/mobile apps, telephony stack, and data processing.
  • Act as a bridge between SRE and software/data engineering. Advocate for operations-focused engineering in an empathetic, contributory spirit. You will treat developers as your customers and build self-service tooling so they can ship without filing tickets.
  • Because our business operates across all US time zones (ET, CT, MT, and PT), you will help stand up the on-call rotation and contribute to shaping sustainable core on-call hours and escalation paths alongside the Director.
  • Lead production incidents with a calm demeanor, drive blameless post-incident reviews (RCAs), and collaborate with teams to fix systemic issues so they don't happen again.
  • Help implement and maintain infrastructure controls (IAM, encryption, network segmentation) that align with HIPAA and HITRUST requirements, ensuring a secure environment for patient data.
  • Other duties as assigned

Qualifications

  • 5+ years in Software Engineering, SRE, DevOps, or Platform Engineering, with a proven track record of operating at a Senior level (building systems, delivering technical initiatives, and collaborating with peers).
  • Deep hands-on AWS expertise (VPC, IAM, ECS/EKS, Lambda, RDS, S3) and production-grade Terraform experience at scale (modules, state management, multi-environment).
  • Strong programming skills in Python, Go, TypeScript, or similar. You treat infrastructure as software and automate away toil.
  • Hands-on experience with modern observability stacks (Datadog, CloudWatch, Grafana, etc.) and a practical understanding of how to implement SLOs, SLIs, and error budgets in a startup environment.
  • Deep experience maintaining and standardizing CI/CD pipelines and tracking delivery metrics (like DORA).
  • Experience operating in a HIPAA, HITRUST, SOC 2 Type II, or comparably regulated growth-stage environment is highly desired.
  • Ability and willingness to travel up to 10% as needed for onsite meetings, team collaboration, and company events.

Skills

  • You own outcomes: When something breaks, you fix it and improve the system so it does not happen again.
  • You write code and ship infrastructure: You are a builder at heart and thrive in hands-on technical environments.
  • You are an empathetic partner: You don't throw policies over the wall; you build paved roads and enable product engineers to do their best work safely.
  • You build for clarity and simplicity: You distrust complexity that does not earn its keep.
  • You bring calm to incidents and discipline to operations.
  • You surface risks early: Bad news early is manageable; bad news late is expensive.
  • Excellent communication skills with a history of building high-trust partnerships with software engineering teams.
  • Experience integrating or operating Salesforce and telephony/contact center platforms (like Amazon Connect) is a plus.
  • Familiarity with data platforms (Databricks, Snowflake, Fivetran) is a plus.
  • Interest or experience in AI/LLM-assisted operations tooling and automation is a plus.

What Success Looks Like

  • Within 30 Days: Gain a deep understanding of our existing AWS footprint, observability capabilities, CI/CD processes, and access controls. Form meaningful, initial partnerships with software and data engineering teams.
  • Within 90 Days: Standardize and modularize critical parts of our Terraform codebase. Roll out the foundational pieces of our universal CI/CD pipeline, audit and upgrade critical observability dashboards to reduce alert fatigue, and participate in standing up a robust infrastructure on-call rotation.
  • In Year One: This role is explicitly hands-on. In your first year, you will personally write production Terraform, review infrastructure pull requests, and pair directly with engineers on critical migrations. You will collaborate with product and engineering teams to support the operational standards expected of the organization and our clients. You will successfully support baseline SLOs, SLIs, and error budgets for TailorCare’s most critical services, establish comprehensive observability (metrics, logs, traces, alerting) across AWS and third-party integrations, and track DORA metrics to ensure reliable, high-velocity delivery.

What's In It For You

  • Meaningful Work: We are dedicated to our mission and deeply value our patients and each other. Each day offers the opportunity to make a positive impact.
  • Work Environment: We operate as a remote-first company with options for a hybrid work model in Nashville.
  • Time Off: Our generous paid time off (PTO) and holiday plans ensure you have ample time to rest and recharge.
  • Family First: We offer paid parental leave and support a healthy work-life balance, encouraging flexibility and autonomy. We love talking about our family and pets!
  • Comprehensive Benefits: From Day 1, employees enjoy medical, dental, vision, life, and disability insurance, wellness resources and an employer HSA contribution.
  • Fair Compensation: We are committed to equitable pay for all team members and support your future goals with a 401k plan that includes employer matching.
  • Community: We foster an inclusive environment where you can rely on your teammates, share honest feedback, and feel comfortable being your authentic self at work each day.

TailorCare seeks to recruit and retain staff from diverse backgrounds and encourages qualified candidates to apply. TailorCare is an equal opportunity employer and does not discriminate on the basis of age, sex, gender identity/expression, sexual orientation, color, race, creed, national origin, ancestry, religion, marital status, political belief, physical or mental disability, pregnancy, military, or veteran status.

Similar roles