Amtech Software

Platform Engineer (SRE - India)- III

Amtech Software Bengaluru, Karnataka, India

Software Development · 51-200 employees

3 h ago
sre Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the reliability of production domains, design observability and automation frameworks, and act as an incident commander for complex issues. Additionally, you will mentor earlier-career engineers and drive global reliability standards across the India and U.S. teams.

What they look for

AWS Terraform Python SRE CI/CD GitHub Actions Kubernetes Observability OpenTelemetry Incident Management PostgreSQL Cloud Security Automation System Reliability Mentorship AI Integration

Requirements

Candidates must have 4-6 years of experience in SRE, DevOps, or cloud engineering with deep knowledge of AWS and Terraform. A bachelor's degree in a related field is required, along with proven ability in Python and experience managing SLO-driven operations.

Full description

SUMMARY

Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.

As a senior individual contributor on the SRE track, you will own the reliability of significant production domains, design the observability and automation frameworks the team standardizes on, and act as incident commander for complex, cross-service incidents.

You will raise the bar for the India SRE pod: setting patterns, reviewing designs, and mentoring earlier-career engineers while remaining deeply hands-on.

JOB DESCRIPTION

Reliability & Performance

  • Design and implement monitoring, alerting, and reliability frameworks on our OpenTelemetry-based stack (OpenObserve, CloudWatch) integrated with PagerDuty.
  • Define and defend SLIs, SLOs, and error budgets for the domains you own, and drive engineering priorities from them.
  • Engineer self-healing, autoscaling, and capacity management so systems recover without human intervention.
  • Lead root cause analysis for high-severity incidents and ensure permanent, verified fixes.

Automation & Operations

  • Design reusable Terraform modules and golden-path patterns adopted across Amtech's multi-account AWS organizations.
  • Build and harden CI/CD pipelines in GitHub Actions, including progressive delivery (blue/green, canary) and automated rollback.
  • Operate and optimize containerized and serverless workloads (ECS Fargate, EKS, Lambda) and RDS PostgreSQL data stores at production scale.
  • Eliminate toil systematically: identify, quantify, and automate the highest-cost operational work.

Incident Response & On-Call

  • Serve in the 24/7 on-call rotation and act as incident commander for complex, multi-service incidents.
  • Own MTTD/MTTR improvement for your domains, with measurable targets.
  • Author runbooks and drive game days or failure testing to validate them.

Security & Compliance

  • Engineer security into the platform: IAM boundaries, secrets management, network segmentation, and policy enforcement.
  • Ensure services meet SOC 2 and ISO 27001 obligations with audit evidence generated automatically where possible.

AI Competency

  • Integrate AI across the delivery workflow with reusable prompts and shared context; mentor earlier-career engineers on effective, safe usage.
  • Design validation steps for AI-assisted changes: characterization tests before AI refactors, empirical verification of AI debugging hypotheses, and rollback plans.
  • Quantify AI productivity impact against baselines rather than impressions, and apply data classification policy to all AI usage.

Technical Leadership & Collaboration

  • Mentor SRE I/II engineers through design review, paired incident response, and code review.
  • Partner with product development and Cloud-track peers on operability and migration readiness for Encore and LabelTraxx workloads.
  • Champion reliability culture and consistent global standards between India and U.S. teams.

QUALIFICATIONS

  • 4-6 years of hands-on SRE, DevOps, or cloud engineering experience, including ownership of production services at meaningful scale.
  • Deep working knowledge of AWS (EC2, ECS/EKS, RDS, S3, IAM, VPC, Lambda) in multi-account environments.
  • Strong Terraform skills, including authoring reusable modules, and fluency with GitHub Actions or equivalent CI/CD.
  • Proven software engineering ability in Python (or similar) applied to automation and tooling, not just scripts.
  • Demonstrated experience running SLO-driven operations, incident command, and blameless post-incident reviews.
  • Experience designing observability with OpenTelemetry or equivalent (metrics, logs, traces).
  • Solid understanding of networking, DNS, and cloud security architecture.
  • Demonstrated integration of AI into daily engineering workflow with measured impact and designed validation, not ad-hoc usage.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.

PREFERRED QUALIFICATIONS

  • AWS Certified DevOps Engineer Professional, CKA, or Terraform Associate.
  • Experience with multi-tenant SaaS or account-per-customer architectures and ERP-class workloads.
  • Experience with PagerDuty at scale (escalation policies, service ownership models).
  • Experience operating AI/LLM workloads or building agentic automation under governance controls.
  • Prior mentorship or tech-lead experience in a distributed global team.

Similar roles