Kong

Site Reliability Engineer 2, Managed Gateways

Kong Bengaluru, Karnataka, India

Software Development · 1,001-5,000 employees

23 h ago
sre Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will be responsible for maintaining the reliability, scalability, and performance of managed services through robust automation and monitoring. Additionally, you will resolve production incidents and collaborate with engineering teams to design resilient, scalable infrastructure.

What they look for

Site Reliability Engineering Golang Python Kubernetes AWS GCP Azure Prometheus Grafana Datadog Networking Distributed systems API gateways Infrastructure as code Automation

Requirements

Candidates must have at least 2 years of experience in Site Reliability Engineering and proficiency in Golang or Python. Hands-on experience with Kubernetes and major cloud platforms like AWS, GCP, or Azure is also required.

Full description

Are you ready to unlock intelligence?

If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

About the Role:

As an SRE 2 for Managed Gateways, you will be pivotal in ensuring the rock-solid reliability, scalability, and performance of Kong's critical managed services. Your expertise will directly impact customer trust and position Kong as a leader in the Agentic Era through unparalleled product stability.

 

What You’ll Do:

  • Implement and maintain robust automation for deploying and operating Kong's Managed Gateways across various cloud environments.
  • Monitor system health, performance, and uptime, striving for 99.99% availability for our core infrastructure.
  • Resolve complex production incidents efficiently, participating actively in on-call rotations to maintain service continuity.
  • Build resilient tools and systems that enhance the overall reliability and operational efficiency of our platform.
  • Contribute proactively to the prevention of technical debt, ensuring sustainable and scalable operations as Kong grows.
  • Collaborate closely with engineering teams to design, review, and implement resilient and highly scalable services.

 

What You’ll Bring:

  • 2+ years of experience applying Site Reliability Engineering (SRE) principles and practices in a production environment.
  • Proficiency in at least one of Golang or Python for automation, tooling, and infrastructure as code.
  • Hands-on experience with Kubernetes and major cloud platforms such as AWS, GCP, or Azure.
  • Familiarity with monitoring, logging, and alerting tools (e.g., Prometheus, Grafana, Datadog).
  • Solid understanding of networking concepts, distributed systems, and API gateways.

 

The Kong DNA:

  • Own the reliability and performance of critical production systems with a strong sense of accountability.
  • Drive urgent resolution of issues, demonstrating a bias for action and minimizing customer impact.
  • Collaborate effectively with cross-functional teams, fostering an environment of shared understanding and collective success.

 

Bonus Points:

  • Experience with Kong Gateway or other API management platforms.
  • Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, Kubernetes Administrator).
  • Active contributions to open-source projects or developer communities.

#LI-AP1

About Kong:

Kong Inc., a leading developer of API and AI connectivity technologies, is building the infrastructure that powers the agentic era. Trusted by the Fortune 500 and startups alike, Kong's unified API and AI platform, Kong Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models. For more information, visit www.konghq.com.

Similar roles