National Payments Corporation of India (NPCI)

Lead DevOps

National Payments Corporation of India (NPCI) · Hyderabad, Telangana, India

Financial Services · 1,001-5,000 employees

3 d ago
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will architect, build, and maintain resilient CI/CD pipelines while owning the reliability, scalability, and security of Kubernetes clusters. Additionally, you will drive infrastructure-as-code practices, establish observability standards, and mentor team members to improve operational efficiency.

What they look for

Kubernetes CI/CD DevOps Infrastructure as Code Linux Python Bash Terraform Ansible Prometheus Grafana Helm GitLab CI GitHub Actions ArgoCD Service Mesh

Requirements

The role requires 6+ years of experience in DevOps or SRE roles with deep expertise in production-grade Kubernetes environments. Candidates must possess strong Linux internals knowledge, proficiency in CI/CD pipeline design, and experience with automation scripting.

Full description

Lead DevOps Engineer Role Overview We're looking for a Lead DevOps Engineer to own, architect, and scale our infrastructure and CI/CD landscape. You'll be the go-to person for pipeline design, Kubernetes operations, and infrastructure reliability — someone who's been in the trenches of production-grade systems and can guide a team while staying hands-on. Experience & Qualifications

  • 6+ years of experience in DevOps, SRE, or Platform Engineering roles.
  • Proven track record of designing and operating production-grade Kubernetes environments at scale.
  • Prior experience in a lead or senior individual contributor role mentoring engineers and driving DevOps best practices (good to have).

Must-Have Skills CI/CD & Pipelines

  • Expert-level proficiency in building, optimizing, and maintaining CI/CD pipelines from scratch.
  • Deep experience with tools such as GitLab CI, GitHub Actions, Jenkins, or ArgoCD.
  • Strong understanding of pipeline security, artifact management, and deployment strategies (blue/green, canary, rolling).

Operating Systems & Fundamentals

  • Deep understanding of Linux internals — processes, memory management, filesystems, networking, and IPC.
  • Solid grasp of OS-level performance tuning, troubleshooting, and kernel parameters.
  • Strong scripting skills (Bash, Python) for automation and tooling.

Kubernetes (Deep & Thorough)

  • Hands-on experience deploying, managing, and troubleshooting Kubernetes in production.
  • Deep knowledge of K8s architecture — control plane, kubelet, etcd, CNI, CSI, scheduler, and admission controllers.
  • Experience with Helm, Kustomize, Operators, CRDs, and policy management.
  • Strong understanding of workload lifecycle, resource management, autoscaling (HPA/VPA/Cluster Autoscaler), and RBAC.
  • Familiarity with service mesh (Istio/Linkerd) and ingress controllers.

Good to Have

  • Experience working in on-premises environments — bare metal or virtualized infrastructure.
  • Understanding of on-prem Kubernetes setup — RKE2, K3s, or similar provisioning tools.
  • Knowledge of on-prem networking — VLANs, load balancing, DNS, firewall rules, and storage backends (Ceph, NFS, iSCSI).
  • Prior lead DevOps experience — driving technical decisions, roadmap planning, and mentoring junior engineers.

What You'll Do

  • Architect, build, and maintain resilient CI/CD pipelines across multiple services and environments.
  • Own the reliability, scalability, and security of our Kubernetes clusters.
  • Collaborate with development teams to improve deployment velocity and infrastructure self-service.
  • Drive infrastructure-as-code practices using Terraform, Ansible, or equivalent.
  • Establish observability standards — logging, monitoring, and tracing (Prometheus, Grafana, ELK, OpenTelemetry).
  • Lead incident response and post-mortem culture; drive SLO/SLI adoption.
  • Evaluate and introduce new tools and practices to improve operational efficiency.
  • Mentor team members and set DevOps standards across the organization.

Nice to Have

  • Experience with GPU workloads or AI/ML infrastructure.
  • Familiarity with container runtimes beyond containerd (CRI-O, etc.).
  • Exposure to GitOps workflows and progressive delivery tooling.