Sr. DevOps Engineer
FOLD HEALTH Pune City Subdistrict, Maharashtra, India
Hospitals and Health Care · 11-50 employees
About the role
The DevOps Engineer will design, build, and maintain scalable cloud infrastructure across AWS and GCP while managing production environments. They will also implement automation, CI/CD pipelines, and robust monitoring solutions to ensure platform reliability and security.
What they look for
Requirements
Candidates must have a bachelor's degree and at least 5 years of hands-on experience in DevOps or Site Reliability Engineering. Proficiency in AWS, Terraform, Docker, Kubernetes, and scripting languages like Python or Bash is required.
Benefits
Full description
Fold Health | Engineering
DevOps Engineer
Job Title
DevOps Engineer
Experience
5 Years +
Location
Pune (Amar Tech Park, Balewadi) · Full-time · In-office
Domain
Healthcare (Preferred)
About Fold Health
Fold Health is building an AI-powered healthcare technology platform designed to support Value-Based Care (VBC) programs. By combining advanced data integration, analytics, and intelligent automation, we help providers, payers, and care teams deliver better outcomes and improve patient experiences. Our mission is to simplify healthcare technology, ensure interoperability, and enable innovation at scale. Join us and be part of shaping the future of AI-driven healthcare.
Role Overview
We are looking for a DevOps Engineer who will be responsible for building, maintaining, and scaling the cloud infrastructure and delivery pipelines that power Fold Health's healthcare platform. The ideal candidate is hands-on, reliability-focused, and comfortable working across AWS, GCP, Terraform, and Prometheus in a fast-paced healthcare environment.
Responsibilities
1. Cloud Infrastructure & Platform Engineering
- Design,
build, and maintain highly available, secure, and scalable cloud infrastructure across AWS and GCP.
- Manage
production environments including Amazon ECS Fargate, RDS PostgreSQL, ElastiCache, ALB, SQS/SNS, Lambda, Step Functions, Cognito, Route 53, and supporting GCP services.
- Optimize
cloud resources for performance, reliability, availability, and cost efficiency.
- Design and
maintain networking, security groups, load balancing, and disaster recovery capabilities.
2. Infrastructure as Code & Automation
- Develop
and maintain reusable Infrastructure as Code using Terraform following industry best practices.
- Build,
enhance, and maintain CI/CD pipelines to enable reliable, secure, and automated application deployments.
- Automate
infrastructure provisioning, operational tasks, and cloud governance through scripting and tooling.
- Standardize
infrastructure components, deployment patterns, and operational workflows across environments.
3. Observability, Monitoring & Incident Management
- Build and
maintain comprehensive monitoring, logging, and alerting solutions using CloudWatch, Prometheus, AlertManager, Grafana, and xMatters.
- Develop
actionable alerting strategies to reduce alert fatigue while ensuring rapid detection of production issues.
- Create
dashboards, metrics, and operational insights to improve platform health and service reliability.
- Participate
in production incident response, root cause analysis, and post-incident reviews, driving preventive improvements.
4. Reliability, Performance & Security
- Ensure
platform reliability, scalability, and performance through proactive capacity planning, tuning, and optimization.
- Troubleshoot
complex production issues across infrastructure, networking, databases, and distributed applications.
- Implement
security best practices including IAM, secrets management, encryption, vulnerability remediation, and least-privilege access.
- Support
compliance initiatives such as HIPAA, SOC 2, and HITRUST by implementing required technical controls and maintaining audit readiness.
5. Collaboration & Continuous Improvement
- Partner
with software engineering teams to improve application deployment, reliability, and operational excellence.
- Advocate
DevOps best practices, automation, and Infrastructure as Code throughout the engineering organization.
- Continuously
evaluate and adopt new cloud technologies, tools, and processes to improve platform efficiency and developer experience.
- Create and
maintain technical documentation, runbooks, and operational procedures to support knowledge sharing and onboarding.
Requirements
Requirements
· Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
· Minimum 5 years of hands-on experience in DevOps, Platform Engineering, or Site Reliability Engineering.
· Strong experience with AWS services such as ECS, RDS, IAM, CloudWatch, ALB, Lambda, and networking.
· Hands-on experience with Infrastructure as Code using Terraform and CI/CD pipelines using GitLab CI/CD or similar tools.
· Strong experience with Docker and containerized application deployments.
· Working knowledge of Kubernetes, including deployments, services, configuration, scaling, and troubleshooting.
· Familiarity with PostgreSQL, Redis or ElastiCache, and microservices architectures.
· Proficiency in Linux administration, Python or Bash scripting, and infrastructure automation.
· Experience with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, and AlertManager.
· Good understanding of cloud security, IAM, networking, secrets management, and production troubleshooting.
· Strong problem-solving, communication, and collaboration skills.
Good to Have
· Experience with Google Cloud Platform services.
· Experience managing Kubernetes workloads in production using EKS, GKE, or similar platforms.
· Experience with on-call management tools such as xMatters or PagerDuty.
· Knowledge of healthcare compliance standards such as HIPAA, HITRUST, or SOC 2.
Benefits
Why Join Us?
- Opportunity to build and scale
an AI-powered healthcare platform impacting millions of lives
- Work with modern cloud
technologies in a mission-driven environment
- Collaborative and innovative
culture with strong growth opportunities
Similar roles
-
DevOps Engineer - Data & Machine Learning Platform (f/m/div.)
Bosch Group Braga, Portugal
-
DevOps Manager, Platform
Origami Risk LLC United States · $145K–$185K/yr
-
IT DevOps Engineer
Entirely AG United States
-
Senior DevOps Engineer
EXL Ciudad de México, Mexico
-
Cloud Systems Engineer - DevOps
TherapyNotes.com Philadelphia, Pennsylvania, United States · $110K–$150K/yr
-
DevOps Engineer
EasySend Tel Aviv, Tel-Aviv District, Israel