Lead DevOps
National Payments Corporation of India (NPCI) · Hyderabad, Telangana, India
Financial Services · 1,001-5,000 employees
About the role
You will architect, build, and maintain resilient CI/CD pipelines while owning the reliability, scalability, and security of Kubernetes clusters. Additionally, you will drive infrastructure-as-code practices, establish observability standards, and mentor team members to improve operational efficiency.
What they look for
Requirements
The role requires 6+ years of experience in DevOps or SRE roles with deep expertise in production-grade Kubernetes environments. Candidates must possess strong Linux internals knowledge, proficiency in CI/CD pipeline design, and experience with automation scripting.
Full description
Lead DevOps Engineer Role Overview We're looking for a Lead DevOps Engineer to own, architect, and scale our infrastructure and CI/CD landscape. You'll be the go-to person for pipeline design, Kubernetes operations, and infrastructure reliability — someone who's been in the trenches of production-grade systems and can guide a team while staying hands-on. Experience & Qualifications
- 6+ years of experience in DevOps, SRE, or Platform Engineering roles.
- Proven track record of designing and operating production-grade Kubernetes environments at scale.
- Prior experience in a lead or senior individual contributor role mentoring engineers and driving DevOps best practices (good to have).
Must-Have Skills CI/CD & Pipelines
- Expert-level proficiency in building, optimizing, and maintaining CI/CD pipelines from scratch.
- Deep experience with tools such as GitLab CI, GitHub Actions, Jenkins, or ArgoCD.
- Strong understanding of pipeline security, artifact management, and deployment strategies (blue/green, canary, rolling).
Operating Systems & Fundamentals
- Deep understanding of Linux internals — processes, memory management, filesystems, networking, and IPC.
- Solid grasp of OS-level performance tuning, troubleshooting, and kernel parameters.
- Strong scripting skills (Bash, Python) for automation and tooling.
Kubernetes (Deep & Thorough)
- Hands-on experience deploying, managing, and troubleshooting Kubernetes in production.
- Deep knowledge of K8s architecture — control plane, kubelet, etcd, CNI, CSI, scheduler, and admission controllers.
- Experience with Helm, Kustomize, Operators, CRDs, and policy management.
- Strong understanding of workload lifecycle, resource management, autoscaling (HPA/VPA/Cluster Autoscaler), and RBAC.
- Familiarity with service mesh (Istio/Linkerd) and ingress controllers.
Good to Have
- Experience working in on-premises environments — bare metal or virtualized infrastructure.
- Understanding of on-prem Kubernetes setup — RKE2, K3s, or similar provisioning tools.
- Knowledge of on-prem networking — VLANs, load balancing, DNS, firewall rules, and storage backends (Ceph, NFS, iSCSI).
- Prior lead DevOps experience — driving technical decisions, roadmap planning, and mentoring junior engineers.
What You'll Do
- Architect, build, and maintain resilient CI/CD pipelines across multiple services and environments.
- Own the reliability, scalability, and security of our Kubernetes clusters.
- Collaborate with development teams to improve deployment velocity and infrastructure self-service.
- Drive infrastructure-as-code practices using Terraform, Ansible, or equivalent.
- Establish observability standards — logging, monitoring, and tracing (Prometheus, Grafana, ELK, OpenTelemetry).
- Lead incident response and post-mortem culture; drive SLO/SLI adoption.
- Evaluate and introduce new tools and practices to improve operational efficiency.
- Mentor team members and set DevOps standards across the organization.
Nice to Have
- Experience with GPU workloads or AI/ML infrastructure.
- Familiarity with container runtimes beyond containerd (CRI-O, etc.).
- Exposure to GitOps workflows and progressive delivery tooling.