DDN

SRE Engineer

DDN · Pune, Maharashtra, India

Software Development · 1,001-5,000 employees

20 h ago
Remote Principal (10+ yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will lead the design and evolution of hybrid infrastructure and developer platforms while ensuring system availability through automation and observability. The role involves collaborating with cross-functional teams to resolve deployment blockers and improve overall platform reliability.

What they look for

Site Reliability Engineering Infrastructure Engineering Cloud Automation Bare-metal Systems DevOps Python Go Terraform Helm Kubernetes Observability Telemetry CI/CD System Architecture Incident Response GPU-accelerated Compute

Requirements

Candidates must have over 10 years of experience in infrastructure or DevOps engineering with strong programming skills in Python or Go. Proficiency in hybrid infrastructure, Kubernetes, and infrastructure-as-code tools like Terraform is required.

Full description

As a Staff Engineer – Site Reliability Engineering, you will play a key technical leadership role in designing, building, and evolving our hybrid infrastructure and developer platforms. You’ll work hands-on across cloud automation, high-performance bare-metal systems, and DevOps tooling to deliver reliable, scalable infrastructure that accelerates software delivery. This role is highly cross-functional. You will collaborate closely with Development, DevOps, and QA teams to solve infrastructure and release challenges, support delivery cycles, and continuously improve platform capabilities. You’ll balance strong technical execution with architectural influence, contributing to system design while remaining directly involved in implementation and operational support. Occasional in-person meetings or team events may be required.

Key Responsibilities

Infrastructure & Platform Engineering

  • Ensure system availability and reliability through automated monitoring strategies
  • Produce post-mortems and implement resulting process improvements
  • Proactively mitigate operational risks through risk assessment and wider collaboration with
  • engineering teams
  • Design, implement, and continuously measure and improve risk mitigation strategies
  • Monitor system health through observability and telemetry
  • Unblock bottlenecks in system performance
  • Minimize emergency response time priods
  • Maintain internal tooling surrounding bug tracking, CI/CD pipelines, and wider
  • communication with the teams

Engineering Collaboration & Delivery Support

  • Partner with Dev, DevOps, and QA teams to resolve infrastructure or deployment blockers during release cycles
  • Provide technical guidance and mentorship to platform engineers
  • Participate in architectural reviews, release readiness checkpoints, and root-cause analyses

Required Qualifications

  • 10+ years of experience in infrastructure, platform, or DevOps engineering roles
  • Strong programming skills (e.g., Python, Go)
  • Hands-on experience with hybrid infrastructure (cloud + bare metal)
  • Deep knowledge of infrastructure-as-code tools (Terraform, Helm, Kubernetes)
  • Proven cross-functional collaboration skillsPreferred Qualifications
  • Experience with GPU-accelerated compute or HPC-style infrastructure
  • Familiarity with platform engineering or developer experience optimization
  • Experience with high-velocity release cycles and incident response

Preferred Qualifications

  • Experience with GPU-accelerated compute or HPC-style infrastructure
  • Familiarity with platform engineering or developer experience optimization
  • Experience with high-velocity release cycles and incident response