Senior Software Engineer - Site Reliability
Synthlane Technologies Private Limited · Gurgaon, Haryana, India
IT Services and IT Consulting · 11-50 employees
About the role
You will lead the design, implementation, and evolution of highly available, scalable, and resilient multi-cloud infrastructure. Additionally, you will drive automation strategies, manage incident responses, and mentor teams to ensure operational excellence.
What they look for
Requirements
Candidates must have 5+ years of experience in Site Reliability Engineering with deep expertise in AWS, Azure, and Infrastructure as Code tools like Terraform. Strong proficiency in Kubernetes, Linux, and scripting languages such as Python or Go is required to manage large-scale distributed systems.
Full description
We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the
design, implementation, and evolution of highly available, scalable, and resilient systems
across our multi-cloud infrastructure. In this senior role, you will drive architectural
decisions, establish reliability standards, and mentor teams while ensuring operational
excellence across complex distributed systems. You will partner with engineering
leadership, development teams, and product stakeholders to shape infrastructure
strategy, implement sophisticated automation, and champion a culture of reliability
engineering. As a Senior SRE, you'll tackle sophisticated, large-scale challenges using cutting-edge
technologies across AWS and Azure platforms. You will lead critical initiatives that
impact system reliability at scale, architect solutions for complex infrastructure problems,
and guide teams in adopting industry-leading practices that drive meaningful
improvements across our entire technology ecosystem.
Requirements
● Architect and implement highly reliable, scalable, and cost-effective infrastructure
solutions for mission-critical applications across multi-cloud environments (AWS and
Azure).
● Lead the definition and refinement of service level objectives (SLOs), service level
indicators (SLIs), and error budgets, establishing reliability standards across the
organization.
● Design and implement sophisticated Infrastructure as Code (IaC) solutions using
Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
● Drive automation strategies to eliminate toil, improve operational efficiency, and enable
self-service capabilities for development teams.
● Lead incident response efforts, conduct thorough post-incident reviews, and implement
systemic improvements to prevent recurrence.
● Champion cloud-native architectures and modern reliability practices, serving as a
technical advisor for infrastructure and platform decisions.
● Participate in and help optimize the on-call rotation, ensuring sustainable practices and
effective escalation procedures.
● Establish and maintain comprehensive documentation standards, runbooks, and
knowledge repositories that enable team autonomy and effective incident response.
● Design and implement advanced monitoring, logging, and alerting strategies using
observability platforms to enable proactive issue detection and resolution.
● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
sophisticated deployment strategies including blue-green, canary, and progressive
delivery patterns.
● Ensure security, compliance, and governance standards are embedded throughout the
infrastructure lifecycle, implementing security-as-code practices.
● Drive capacity planning, performance optimization, and cost management initiatives
across cloud platforms.
● Collaborate with architecture and security teams to establish platform standards,
reference architectures, and best practices.
Skills, Knowledge and Expertise
● 5+ years of proven experience as a Site Reliability Engineer or similar role, with
demonstrated expertise in designing, implementing, and operating large-scale,
distributed systems.
● Deep expertise in Infrastructure as Code (IaC) with Terraform and Ansible, including
module development, state management, and multi-environment orchestration.
● Extensive hands-on experience with both AWS and Azure cloud platforms, including
advanced services, networking, and security features in both environments.
● Expert-level knowledge of container orchestration with Kubernetes, including
architecture, custom resource definitions (CRDs), operators, service mesh
implementations, and production-scale cluster management.
● Advanced proficiency in Linux system administration, performance tuning, and
troubleshooting complex system-level issues.
● Proven experience implementing GitOps workflows using ArgoCD, Flux, or similar tools,
including advanced deployment patterns and progressive delivery.
● Deep understanding of observability principles and hands-on experience with tools such
as Prometheus, Grafana, Datadog, Azure Monitor, or the ELK stack.
● Expert knowledge of networking concepts, including load balancing, CDNs, DNS, VPNs,
service mesh architectures, and distributed systems communication patterns.
● Strong programming and scripting capabilities in Python, Bash, Go, or PowerShell, with
the ability to develop custom tooling and automation frameworks.
● Extensive experience designing and optimizing CI/CD pipelines using Jenkins, GitLab
CI, Azure DevOps, GitHub Actions, or CircleCI.
● Demonstrated ability to lead incident response, conduct root cause analysis, and drive
systemic reliability improvements.
● Excellent communication and leadership skills with proven ability to influence technical
decisions and collaborate with stakeholders at all levels.
● Current certification in AWS (Solutions Architect Associate/Professional or equivalent)
and Azure (Azure Administrator or Azure Solutions Architect), with practical experience
managing production workloads on both platforms.
Good To Have
● Experience with hybrid and multi-cloud networking strategies, including ExpressRoute,
Direct Connect, and cloud interconnects.
● Knowledge of serverless architectures on AWS (Lambda) and Azure (Functions, Logic
Apps) and their operational considerations.
● Proven experience with disaster recovery planning, business continuity, and
implementing multi-region active-active architectures.
● Understanding of machine learning operations (MLOps), data pipeline orchestration, and
supporting ML workloads in production.
● Experience with service mesh technologies such as Istio, Linkerd, or Consul.
● Familiarity with chaos engineering principles and tools like Chaos Monkey or Gremlin.
● Experience with configuration management at scale and policy-as-code tools like Open
Policy Agent (OPA).
● Knowledge of FinOps principles and cloud cost optimization strategies.
Benefits
Why Work at Synthlane?
Real production infrastructure exposure (not just toy tasks). Work closely with engineers building real solutions and client deployments. Learn what “high reliability” actually means in production. Fast-paced, high-impact environment where good work is noticed quickly.