Senior Site Reliability Engineer — CDN Infrastructure
VergeCloud Bengaluru, Karnataka, India
Technology, Information and Internet · 11-50 employees
About the role
You will operate and maintain a fleet of baremetal edge servers and manage the Kubernetes-based control plane. Additionally, you will build CI/CD pipelines and maintain observability tooling to ensure network performance and reliability.
What they look for
Requirements
Candidates must have 5+ years of experience in SRE or infrastructure engineering with a deep understanding of networking fundamentals like DNS and TCP/IP. Proficiency in Linux, configuration management tools, and automation scripting in Go or Python is required.
Full description
About the Role
We're looking for a Senior SRE with 5+ years of experience to help operate and evolve
the infrastructure behind our CDN — from the baremetal edge nodes serving traffic, to
the Kubernetes-based control plane, to the observability and data pipelines that tell us
what's actually happening on the network. You'll work across the full stack: provisioning
and configuration management, CI/CD, Kubernetes, and the logging/monitoring systems
that let us catch problems before customers do.
This is a hands-on, senior individual-contributor role for someone who's comfortable
moving between low-level networking debugging and higher-level platform/infra
automation, and who can drive decisions with less oversight.
What You'll Do
Operate and maintain our fleet of self-managed baremetal edge servers, using
SaltStack (or Ansible) for configuration management and automation
Manage and improve our managed Kubernetes cluster, deploying and maintaining
services via Helm
Build and maintain CI/CD pipelines (GitHub Actions) for infrastructure and service
deployments
Own and extend observability tooling — Prometheus, Grafana, Alertmanager, and
distributed tracing with Grafana Tempo
Maintain and query ClickHouse at scale — millions of rows ingested daily, with a
focus on schema design and low-latency query tuning
Manage cloud service configuration as code using Terraform
Diagnose and resolve issues across the stack — from DNS resolution and
BGP/routing anomalies to TCP/IP-level performance regressions and applicationWhat
You'll Need
Requirements
What You'll Need
5+ years of experience in an SRE, DevOps, or infrastructure/platform engineering
role
Deep understanding of networking fundamentals — DNS, TCP/IP, and CDN concepts
(caching, routing, anycast, edge delivery) — you should be comfortable reading a
packet capture or debugging a DNS resolution chain
Solid experience operating Linux systems in production, including self-managed
baremetal infrastructure
Hands-on experience with configuration management tools — SaltStack or Ansible
Experience with Kubernetes in production, including deploying and managing
services via Helm
Experience building CI/CD pipelines, ideally with GitHub Actions
Working knowledge of Terraform or similar IaC tools
Practical experience with Prometheus, Grafana, and Alertmanager for monitoring and
alerting; familiarity with distributed tracing (Grafana Tempo or similar)
Ability to write code in Go or Python for automation, tooling, or internal services
Strong debugging skills across layers — from kernel/network to application to
infrastructure automation
Good communication skills and comfort working in a small, high-ownership team
Nice to Have
Experience with BGP / anycast routing in a production CDN or network operator
context
Experience with ClickHouse or another columnar/analytical database at scale
Experience with Kafka/Redpanda or similar streaming systems for log/data pipelines
Prior experience at a CDN, ISP, hosting provider, or similar network-heavy operator
Experience with GitOps workflows (ArgoCD or similar)
Similar roles
-
Site Reliability Engineer II
Backblaze External Website Bangalore, Karnataka, India
-
Site Reliability Engineer
Firmus Technologies Singapore, Singapore
-
Site Reliability Engineer I
Backblaze External Website Bangalore, Karnataka, India
-
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic San Francisco, California, United States · $405K–$485K/yr
-
Linux Site Reliability Engineer (SRE)
OCBC Sepang, Selangor, Malaysia
-
Senior Lead Site Reliability Engineer
FIS pune, Maharashtra, India