Senior Site Reliability Engineer, CloudOps
ICU Medical United States
Medical Equipment Manufacturing · 5,001-10,000 employees
About the role
The Senior Site Reliability Engineer will manage and optimize a multi-account AWS environment while supporting production deployments and infrastructure automation. They will also participate in incident response, on-call rotations, and drive the transition toward containerized and Kubernetes-based platforms.
What they look for
Requirements
Candidates must have at least 7 years of hands-on experience in AWS Cloud Engineering, DevOps, or SRE, along with a Bachelor's degree. Proven experience in regulated environments like HIPAA/HiTrust and strong proficiency in Linux, Python, and AWS core services is required.
Full description
Position Summary
We are seeking a Senior Site Reliability Engineer, CloudOps to support, scale, and optimize a multi-account AWS environment hosting healthcare-oriented applications and analytics platforms. In this role, you will bridge infrastructure engineering, operational reliability, production support, and cloud modernization initiatives across complex microservices architectures. The ideal candidate brings strong AWS expertise, solid Linux administration skills, and a proven track record of managing production systems in HIPAA/HiTrust regulated environments. You will participate in incident response, on-call rotations, and continuous deployment workflows while helping drive our transition toward containerized and Kubernetes-based platforms. This collaborative position is built for an analytical engineer who excels at resolving production incidents, partnering with developers, and continuously elevating operational excellence.
Essential Duties & Responsibilities
- Manage, maintain, and troubleshoot a multi-account AWS Organization environment (35+ accounts) and core services, including EC2, ECS/Fargate, Lambda, S3, CloudFront, API Gateway, and Aurora/RDS databases.
- Support production deployments, CI/CD pipelines (Jenkins, AWS CodePipeline), and infrastructure automation using Python, Bash, and AWS CloudFormation.
- Monitor system health and performance using Datadog, CloudWatch, and Zabbix; investigate alerts, execute root-cause analysis, and refine monitoring coverage to reduce operational noise.
- Participate in a shared on-call rotation, managing incident response and performing failover/recovery validation for production applications and data stores.
- Maintain HIPAA/HiTrust compliance and security posture by managing tools like Prisma/Cortex Cloud, Security Hub, and GuardDuty, while enforcing proper IAM policies and network segmentation.
- Support Java (Spring Boot) and Python applications running in containers, assisting developers during investigations and preparing for future Kubernetes (EKS) modernization initiatives.
Knowledge & Skills
- Deep hands-on expertise with AWS core services (networking, compute, serverless, and database technologies) and CloudFormation IaC automation.
- Strong Linux administration skills (primarily Ubuntu) along with proficiency in Python and Bash scripting for operational automation.
- Experience with containerization technologies (Docker, ECS/Fargate) and familiarity with modern Kubernetes ecosystems (EKS, Helm, ArgoCD).
- Solid understanding of observability tools (Datadog, CloudWatch, Zabbix) and CI/CD pipelines (Jenkins, CodePipeline, Git workflows).
- Knowledge of cloud security best practices, access management (IAM), and compliance frameworks within regulated sectors (HIPAA/HiTrust).
- Proven diagnostic, incident-management, and analytical troubleshooting skills for complex microservices architectures.
Minimum Qualifications, Education & Experience
- Must be at least 18 years of age.
- High School Diploma required.
- Bachelor’s degree from an accredited college or university is required.
- 7+ years of hands-on experience in AWS Cloud Engineering, DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
- Practical background supporting production workloads in Linux/AWS environments, reading application logs, and making minor code fixes.
- Direct experience participating in on-call rotations and incident response protocols.
- Prior experience in the healthcare industry maintaining HIPAA/HiTrust-compliant infrastructure.
Work Environment
- This is largely a sedentary role.
- This job operates in a professional office environment and routinely uses standard office equipment.
- Typically requires travel less than 5% of the time
ICU Medical has consistently provided you with clinical innovations that help solve real-world challenges.
With the acquisition of Hospira Infusion Systems in 2017 and Smiths Medical in 2022, we are now a global market leader with a complete line of clinically-essential IV therapy and high-value critical care products for hospital, alternate site, and home care settings.
We're ready to bring you consistent quality, innovation, and value in more areas than ever. Our focus allows us to bring you:
- Dedicated and non-dedicated IV sets and needlefree connectors clinically proven to provide an effective barrier against bacterial transfer and colonization.
- The industry’s broadest IV smart pump offering covering large volume, pain management, and ambulatory needs.
- IV medication safety software providing full IV-EHR interoperability with the highest customer satisfaction and compatibility with more EHR systems than any other company.
- Significant US IV solutions manufacturing and supply capabilities.
This role is based remotely; the incumbent may be remote in any state other than Colorado; California; Connecticut; Montana, Maine or New York.
ICU Medical EEO Statement:
ICU Medical is committed to being an Equal Opportunity Employer. We ensure that all qualified applicants receive fair consideration for employment regardless of race, color, nationality or national origin, ethnicity, sex, gender, religion or belief, marital or civil partnership status, sexual orientation, pregnancy or maternity, age, disability, or protected veteran status.
If you are an individual with a disability and need reasonable accommodation to participate in the employment selection process, please contact us at humanresources@icumed.com. We are committed to providing equal access and opportunities for all candidates.
ICU Medical EEO Policy Statement
Know Your Rights: Workplace Discrimination is Illegal Poster
ICU Medical CCPA Notice to Job Applicants
Similar roles
-
Site Reliability Engineer III
JPMorgan Chase & Co. Jersey City, New Jersey, United States · $138K–$185K/yr
-
Site Reliability Engineer (Manufacturing Infrastructure)
SpaceX Bastrop, Texas, United States
-
Site Reliability Engineer (SRE) - Engineering Productivity
Jobgether India
-
IC3 - Infra Engineer - SRE
Spin Careers Ciudad de México, Mexico
-
Engineering Manager, Site Reliability
Booking Holdings Bangalore South, Karnataka, India
-
Senior Site Reliability Engineer
ServiceTitan California, United States · $138K–$221K/yr