DevOps / Site Reliability Engineer – Engineering
Practice By Numbers Kolkata, West Bengal, India
Software Development · 51-200 employees
About the role
The engineer will own and extend Terraform modules across AWS infrastructure and build robust CI/CD pipelines for service deployment. They are responsible for maintaining production reliability, managing database migrations, and improving observability through proactive monitoring and automation.
What they look for
Requirements
Candidates must have 6+ years of experience in DevOps or SRE roles with strong proficiency in AWS, Terraform, and container orchestration. A solid background in Linux, Python scripting, and PostgreSQL administration is required, along with the ability to work in either normal or night shifts.
Full description
DevOps / Site Reliability Engineer – Engineering
Experience: 6+ Years
Location: Kolkata, India / Gurgaon, India (On-site)
Shift: Normal or Night — depending on candidate preference and business need Reports to: Lead DevOps / SRE
Employment Type: Full-time
About Practice by Numbers
Practice by Numbers (PbN) is a fast-growing SaaS platform helping healthcare organizations leverage data, automation, patient engagement, and operational intelligence to improve business performance and patient outcomes. We build highly scalable cloud-native products serving thousands of customers across North America.
We are looking for a DevOps/SRE Engineer who writes Terraform and CI/CD pipelines. This is a builder’s role, not a monitoring seat.
Role Overview
You will own Terraform modules across our AWS infrastructure — compute, networking, databases, IAM, DNS — and build the GitHub Actions pipelines that ship our Django, Python, Node.js, and React services.
This role can be staffed on a normal shift or a night shift, depending on your preference and our coverage needs. On a night shift you get an uninterrupted window for the changes that are hardest to make in daylight — database migrations, pipeline rewrites, provider upgrades, and staged cutovers — and you will often be the only engineer on shift. That means real autonomy, real ownership, and the expectation that you can reason through an unfamiliar system on your own and write up clearly what you did.
This role suits someone who wants deep infrastructure work with production ownership and prefers building over ticket-shuffling.
Key Responsibilities
Infrastructure as Code
- Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch.
- Manage Terraform state safely; write and review plans that reviewers can trust.
- Eliminate manually created resources by bringing them under code (via terraform import, refactors, and module extraction).
- Keep environments (dev, QA, production) consistent and reproducible.
CI/CD Engineering
- Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment.
- Migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single, maintainable platform.
- Design safe deployment paths — staged rollouts, health-gated releases, fast and reliable rollback. Keep pipelines fast: caching, parallelism, test splitting, and honest gating.
Reliability & Operations
- Own the production change window on your shift: deploys, migrations, cutovers, and maintenance.
- Run and improve observability — dashboards, SLOs, alerting that fires on real user impact rather than noise.
- Debug containerized applications in production: logs, metrics, traces, resource limits, networking.
- Participate in on-call rotation and incident response; drive blameless post-incident reviews. Improve cost efficiency and resource utilization across the AWS footprint.
Databases & Data Services
- Operate PostgreSQL in production — read query plans, identify slow queries, understand connections, locks, and replication.
- Plan and execute schema migrations against live systems with minimal disruption.
- Operate supporting data services: Redis/ElastiCache, message brokers, object storage.
Automation & Tooling
- Write Python and shell automation to remove repetitive operational work.
- Build tooling that makes the wider team faster — self-service scripts, runbooks, guardrails.
- Harden secrets handling, access control, and infrastructure security posture.
Handover & Communication
- Write clear, complete handover notes at the end of every shift. Where shifts do not overlap, your writing is how the rest of the team learns what happened.
- Maintain runbooks and infrastructure documentation as systems change.
- Coordinate with the wider engineering team on planned work and follow-ups.
AI-Enabled Engineering
- Use modern AI tools to accelerate infrastructure work, scripting, and troubleshooting.
- Apply AI-assisted practices while maintaining strong engineering, security, and review standards.
Required Qualifications
- 6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership.
- Strong AWS experience — container orchestration (ECS/Fargate or EKS), RDS, VPC networking, IAM, load balancing, S3, CloudWatch.
- Hands-on Terraform: writing modules, managing state, reviewing plans.
- Practical CI/CD experience, ideally GitHub Actions. GitLab CI, CircleCI, or Jenkins backgrounds are fine if you can migrate.
- Docker, and comfort debugging containerized applications in production.
- Solid Linux fundamentals and shell scripting; Python for automation.
- Working PostgreSQL knowledge — query plans, slow queries, connections, locks, replication. Experience deploying and operating Django/Python and Node.js/React applications. Hands-on use of an observability platform in anger (New Relic, Datadog, Grafana, Prometheus, or similar).
- Clear written English. Your handover notes are how the rest of the team learns what happened. Openness to either a normal or a night shift, and willingness to participate in an on-call rotation.
Technical Expertise
Cloud & Infrastructure
- AWS (ECS / Fargate / EKS, EC2, Lambda)
- VPC, Subnets, Security Groups, NAT, Peering
- ALB / NLB, Route 53, CloudFront
- IAM, Roles, Policies, Least-Privilege Design
- RDS, ElastiCache, S3
- Terraform / Infrastructure as Code
CI/CD & Automation
- GitHub Actions
- Jenkins / CircleCI / GitLab CI
- Docker, Container Registries
- Blue-Green & Staged Deployments, Rollback Strategies
- Python, Bash / Shell Scripting
- Git and Trunk-Based Workflows
Observability & Reliability
- New Relic / Datadog / Grafana / Prometheus
- CloudWatch Metrics, Logs, Alarms
- OpenTelemetry
- SLOs, Error Budgets, Alert Design
- Incident Management & Post-Incident Review
Data & Messaging
- PostgreSQL Administration & Tuning
- Schema Migrations on Live Systems
- Redis / ElastiCache
- Kafka / Redpanda / Amazon MSK, RabbitMQ, NATS
Security & Compliance
- Secrets Management (AWS Secrets Manager, SOPS, Vault)
- Network and Access Hardening
- Vulnerability and Patch Management
- Audit Logging
Preferred Qualifications
- Celery, RabbitMQ, NATS, or Kafka in production.
- Redis / ElastiCache operations.
- Secrets management (AWS Secrets Manager, SOPS, Vault).
- Experience with sharded or multi-tenant database architectures.
- Healthcare or other compliance-sensitive environments (HIPAA, SOC 2).
- Telephony / VoIP infrastructure exposure.
- Kubernetes.
- Prior experience as the sole engineer on shift, or on a night/off-hours rotation.
What We Look For
- A builder’s instinct — you would rather codify a fix than repeat it.
- Comfort with autonomy and sound judgment on when to escalate.
- Careful, methodical change management on production systems.
- Strong written communication; your handover notes are a first-class deliverable.
- Curiosity about how systems actually behave, not just how they are supposed to. Ownership, follow-through, and continuous learning.
Why This Role
- Real ownership of infrastructure, not a ticket queue.
- Flexibility on shift, and a protected change window for high-impact work.
- Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, Kafka-compatible messaging.
- Direct impact on the reliability of a platform used by thousands of healthcare practices.
Similar roles
-
Senior Software Engineer - SRE & AIOps
ServiceNow Santa Clara, California, United States · $143K–$243K/yr
-
Senior Staff Software Engineer – SRE & AIOps
ServiceNow Santa Clara, California, United States · $191K–$334K/yr
-
Site Reliability Engineer Manager I
Stone - Linkedin Hamburg, Germany
-
ASKUSR0145930 Site Reliability Engineer (SRE)
Essnova Solutions, Inc. Berkeley, California, United States · $166K/yr
-
Head of Site Reliability Engineering (SRE)
Computershare Bristol, England, United Kingdom
-
Site Reliability Engineer, Infrastructure Engineering
CoreWeave Europe Warsaw, Masovian Voivodeship, Poland · PLN 223K–PLN 298K/yr