Razorpay Software Private Limited

Site Reliability Engineer

Razorpay Software Private Limited Bengaluru, Karnataka, India

Software Development · 1,001-5,000 employees

5 h ago
sre Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Build and operate reliable, high-throughput payment infrastructure, including payment processing, payouts, reconciliation, rate limiting, sharding, and traffic management. Improve resilience and operational efficiency through automation, AI-enabled observability and incident response, capacity planning, on-call support, and mentoring teams in SRE practices.

What they look for

Site Reliability Engineering Distributed Systems Software Design and Architecture Go Python Java Automation Capacity Planning Incident Response Observability Machine Learning AIOps Kubernetes Infrastructure as Code Chaos Engineering Stakeholder Management

Requirements

Requires a bachelor's degree or equivalent practical experience, at least five years in software, systems, or site reliability engineering, including three years focused on SRE and three years in software design and architecture. Candidates should be proficient in a programming language such as Go, Python, or Java and have experience with distributed systems and AI-assisted development or operations tools; a master's degree and experience with large-scale systems, AIOps, cloud-native infrastructure, and ML/AI workloads are preferred.

Full description

Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution.

Razorpay powers millions of businesses with a smarter, scalable stack that goes beyond transactions to help them truly build and grow.

From building AI-native agentic payments, to AI-assisted fraud detection and real-time risk intelligence to automated reconciliation, smart payouts, and predictive financial insights, we are embedding intelligence across our stack to make money movement faster, safer, and more efficient. In close collaboration with ecosystem partners - including banks, networks, regulators - we are pioneering industry-first solutions that are shaping the next era of fintech

Across India, Singapore and Malaysia, our products span everything from seamless checkouts to payroll automation - powering a fintech ecosystem that’s redefining how money moves across Asia.

Today, that ecosystem supports everyone from early-stage startups to some of India’s largest enterprises, enabling them to accept, process, and disburse payments at scale while expanding into new ways of managing money more efficiently.

Our scale speaks volumes: Razorpay processes $180+ billion in annualized transactions, powering leading businesses like Airbnb, Facebook, WhatsApp, Airtel, CRED, BookmyShow, Zomato, Swiggy, Lenskart, Mirae Asset Capital markets, Indian Oil, National Pension Scheme - and over 100 of India’s unicorns. With strong roots in India and growing operations in Southeast Asia, we are shaping the next chapter of financial technology across the region.

We are backed by global investors including GIC, Peak XV Partners (formerly Sequoia Capital India & SEA), Tiger Global, Ribbit Capital, Matrix Partners, MasterCard, and Salesforce Ventures, having raised over $740 million to date. Strategic acquisitions - including Ezetap (POS and offline payments), Curlec (Malaysia expansion), BillMe (digital invoicing), and POP (rewards-first UPI) - along with earlier moves in fraud prevention, payroll, and lending, have further strengthened our platform and widened our footprint across Asia.

But what truly sets Razorpay apart is our culture. At Razorpay, ownership is our oxygen - you own what you build, with no micromanagement or red tape, just the runway to make your ideas fly. Learning is a lifestyle - if you’re curious, you’ll feel at home here. People > Pedigree - we hire for attitude, hustle, and hunger more than degrees. Transparency thrives over titles - this is where interns question CXOs and CXOs say “thank you.” Guided by our values of Customer First, Autonomy & Ownership, Agility with Integrity, Transparency, Challenging the status quo and a strong belief that Razorpay grows with Razors, you’ll be part of a 3000+ strong team building not just products, but the financial infrastructure of the future.

About the Role Site Reliability Engineering (SRE) at Razorpay combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems that power India’s digital payment infrastructure. SRE ensures that Razorpay’s services — both our internally critical and our externally- visible systems — have the reliability, uptime, and performance appropriate to the needs of millions of businesses and their customers. Given that every transaction on our platform represents real money movement, the stakes for reliability are exceptionally high. As an SRE, you will keep an ever-watchful eye on system capacity, performance, and latency across a stack processing $180+ billion in annualized transactions. Much of our software development focuses on optimizing existing systems, building infrastructure, and eliminating manual work through automation. On the SRE team, you’ll have the opportunity to manage the complex challenges of scale that are unique to a high-throughput fintech platform, while using your expertise in coding, algorithms, complexity analysis, and large-scale system design. You will increasingly leverage AI and machine learning to transform how we operate — from AI- assisted incident detection and response, to ML driven anomaly detection, predictive capacity planning, and intelligent alerting that reduces noise and surfaces real issues before they impact customers. SRE’s culture of intellectual curiosity, problem solving, and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame free environment.

The SRE team at Razorpay supports services that are foundational to running production at scale — including payment routing, rate limiting, shard management, traffic shaping, and the reliability of core payment, payout, and reconciliation systems. These platforms implement safety guardrails, risk controls, regulatory compliance mechanisms, accounting integrity, and dynamic load balancing for stateful services. Behind everything our customers see is the architecture built by the Infrastructure team to keep it running, and we are proud to be our engineers’ engineers.

Responsibilities

  • Develop strong, influential relationships with multiple stakeholders across the Site Reliability Engineering, Developer, and Product organizations.
  • Serve as an expert on particular fields of knowledge related to rate limiting, sharding, traffic management, and large-scale distributed system reliability for Razorpay’s payment platform.
  • Develop plans and lead projects that evolve our production systems and their reliability — spanning capacity planning, failure-mode analysis, and architectural improvements for high- throughput, low-latency payment infrastructure.
  • Design, build, and operate automation tooling and self-healing infrastructure that eliminates manual toil and improves system resilience across the payments stack.
  • Leverage AI and ML to build intelligent observability — anomaly detection, predictive alerting, noise reduction, and correlation engines that surface real issues before they impact customers.
  • Drive AI-assisted incident management: use LLM-based tooling for root-cause analysis, incident summarization, runbook generation, and post-incident learning to reduce mean-time-to- resolution (MTTR).
  • Build and maintain reliability for AI-native and AI-assisted products — including agentic payments, AI-driven fraud detection, and real-time risk intelligence — ensuring the infrastructure supporting ML models and inference pipelines is production-grade.
  • Identify internal and external opportunities to improve systems, including evaluating emerging AI/Ops (AIOps) tools, observability platforms, and reliability patterns.
  • Provide on-call support for onboarded services, participating in incident response, blameless postmortems, and driving follow-up remediation.
  • Mentor engineers across the organization on SRE practices, reliability thinking, and the effective use of AI tooling in day-to-day operations.
  • Champion SRE best practices — service-level objectives (SLOs), error budgets, toil reduction, and data-driven reliability decisions — across engineering teams.

Minimum Qualifications

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • 5 years of experience in software engineering, systems engineering, or site reliability engineering. 3 years of experience with site reliability engineering focused on building and maintaining scalable, reliable systems.
  • 3 years of experience in software design and architecture, including distributed systems and backend services.
  • Proficiency in at least one programming language (Go, Python, Java, or similar) with the ability to write production-quality code and automation.
  • Working knowledge of AI-assisted development and operations tooling — including experience using LLM-based copilots for code generation, debugging, and incident response workflows.

Preferred Qualifications

  • Master’s degree in Computer Science or a related technical field.
  • 5+ years of experience in large-scale distributed systems, preferably in a high-throughput, mission-critical domain (payments, fintech, e-commerce, or cloud infrastructure).
  • Experience in a Site Reliability Engineering role at scale, with demonstrated impact on system availability, latency, or operational efficiency.
  • Experience leading complex process improvement projects and influencing and managing stakeholder relationships across engineering and product.
  • Experience in troubleshooting and debugging complex distributed systems under production pressure.
  • Hands-on experience with AIOps, ML-based anomaly detection, or building intelligent alerting / observability platforms.
  • Experience with LLM-powered tooling for operations — such as AI-driven incident response, automated root-cause analysis, or intelligent runbook automation.
  • Familiarity with building and operating infrastructure for ML/AI workloads — including model serving, inference pipelines, and the reliability of AI-native products.
  • Knowledge of chaos engineering, fault injection, and resilience testing for distributed systems.
  • Experience with cloud-native infrastructure (Kubernetes, service mesh, cloud platforms) and infrastructure-as-code.

AI Competency At Razorpay, AI is not a bolt-on — it is embedded across our stack and across how we operate. We are building AI-native agentic payments, AI-assisted fraud detection, and real-time risk intelligence. We expect SREs to be fluent in leveraging AI to operate more intelligently, and to build the reliability foundations for AI-powered products. The following AI competencies are expected for this role:

AI for Operations (AIOps)

Experience using LLM-based tools (e.g., coding copilots, AI assistants) in day-to-day SRE work — for code generation, debugging, runbook generation, and incident triage.

  • Ability to build or integrate ML-driven anomaly detection and intelligent alerting that reduces alert noise and surfaces genuine issues before customer impact.
  • Experience with AI-assisted incident response — using LLMs for log correlation, root-cause hypothesis generation, incident summarization, and automated postmortem drafts.
  • Familiarity with predictive capacity planning and forecasting using ML models to anticipate resource needs and prevent bottlenecks.

Reliability for AI Systems

  • Experience building and operating reliable infrastructure for ML/AI workloads — including model serving, inference pipelines, feature stores, and the data plumbing that feeds them.
  • Understanding of the unique failure modes of AI systems — model drift, data quality degradation, inference latency spikes, and hallucination / output reliability — and how to build guardrails and monitoring for them.
  • Ability to define SLOs and error budgets for AI-native products, including agentic payment flows and AI-driven risk decisioning, where the definition of “correctness” is more nuanced than

traditional services. AI-Native Mindset

  • Strong opinion on where AI helps versus where it adds risk in production operations — knowing when to trust AI-driven decisions and when to keep a human in the loop.
  • Eagerness to experiment with emerging AI tooling (agentic frameworks, LLM-based ops assistants, autonomous remediation) and evaluate their applicability to Razorpay’s reliability challenges.
  • Ability to mentor teammates on effective and responsible use of AI in engineering workflows, including prompt hygiene, output verification, and avoiding over-reliance on automated decisions in safety-critical paths.

What You’ll Work On

  • Reliability of Razorpay’s core payment processing, payout, and reconciliation systems — the infrastructure powering $180+ billion in annualized transactions.
  • AI-native and AI-assisted products — including agentic payments, AI-driven fraud detection, and real-time risk intelligence platforms — ensuring they are production-grade and trustworthy.
  • High-throughput, low-latency infrastructure: rate limiting, sharding, traffic management, and dynamic load balancing for stateful services.
  • Intelligent observability: building the next generation of monitoring, alerting, and incident management with AI woven in to reduce noise and improve signal.

Automation and toil elimination: building self-healing systems that detect, diagnose, and remediate issues without human intervention where safe.

  • Capacity and performance: ensuring the platform can scale to support Razorpay’s growing transaction volume and expanding footprint across India and Southeast Asia.

Equal Opportunity Razorpay is an equal opportunity employer. We celebrate diversity and are committed to building an inclusive environment for all Razors. We hire for attitude, hustle, and hunger over mere degrees, and we believe the best teams are made of people with different backgrounds, experiences, and perspectives. If you’re excited about building the financial infrastructure of the future, we’d love to hear from you

Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe.

Follow us on LinkedIn & Twitter

Similar roles