Synthlane Technologies Private Limited

Senior Software Engineer - Site Reliability

Synthlane Technologies Private Limited · Gurgaon, Haryana, India

IT Services and IT Consulting · 11-50 employees

19 h ago
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will lead the design, implementation, and evolution of highly available, scalable, and resilient multi-cloud infrastructure. Additionally, you will drive automation strategies, manage incident responses, and mentor teams to ensure operational excellence.

What they look for

Site Reliability Engineering AWS Azure Terraform Ansible Kubernetes Infrastructure as Code Linux Python Go CI/CD Observability GitOps Networking Distributed Systems Automation

Requirements

Candidates must have 5+ years of experience in Site Reliability Engineering with deep expertise in AWS, Azure, and Infrastructure as Code tools like Terraform. Strong proficiency in Kubernetes, Linux, and scripting languages such as Python or Go is required to manage large-scale distributed systems.

Full description

We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the

design, implementation, and evolution of highly available, scalable, and resilient systems

across our multi-cloud infrastructure. In this senior role, you will drive architectural

decisions, establish reliability standards, and mentor teams while ensuring operational

excellence across complex distributed systems. You will partner with engineering

leadership, development teams, and product stakeholders to shape infrastructure

strategy, implement sophisticated automation, and champion a culture of reliability

engineering. As a Senior SRE, you'll tackle sophisticated, large-scale challenges using cutting-edge

technologies across AWS and Azure platforms. You will lead critical initiatives that

impact system reliability at scale, architect solutions for complex infrastructure problems,

and guide teams in adopting industry-leading practices that drive meaningful

improvements across our entire technology ecosystem.

Requirements

● Architect and implement highly reliable, scalable, and cost-effective infrastructure

solutions for mission-critical applications across multi-cloud environments (AWS and

Azure).

● Lead the definition and refinement of service level objectives (SLOs), service level

indicators (SLIs), and error budgets, establishing reliability standards across the

organization.

● Design and implement sophisticated Infrastructure as Code (IaC) solutions using

Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.

● Drive automation strategies to eliminate toil, improve operational efficiency, and enable

self-service capabilities for development teams.

● Lead incident response efforts, conduct thorough post-incident reviews, and implement

systemic improvements to prevent recurrence.

● Champion cloud-native architectures and modern reliability practices, serving as a

technical advisor for infrastructure and platform decisions.

● Participate in and help optimize the on-call rotation, ensuring sustainable practices and

effective escalation procedures.

● Establish and maintain comprehensive documentation standards, runbooks, and

knowledge repositories that enable team autonomy and effective incident response.

● Design and implement advanced monitoring, logging, and alerting strategies using

observability platforms to enable proactive issue detection and resolution.

● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement

sophisticated deployment strategies including blue-green, canary, and progressive

delivery patterns.

● Ensure security, compliance, and governance standards are embedded throughout the

infrastructure lifecycle, implementing security-as-code practices.

● Drive capacity planning, performance optimization, and cost management initiatives

across cloud platforms.

● Collaborate with architecture and security teams to establish platform standards,

reference architectures, and best practices.

Skills, Knowledge and Expertise

● 5+ years of proven experience as a Site Reliability Engineer or similar role, with

demonstrated expertise in designing, implementing, and operating large-scale,

distributed systems.

● Deep expertise in Infrastructure as Code (IaC) with Terraform and Ansible, including

module development, state management, and multi-environment orchestration.

● Extensive hands-on experience with both AWS and Azure cloud platforms, including

advanced services, networking, and security features in both environments.

● Expert-level knowledge of container orchestration with Kubernetes, including

architecture, custom resource definitions (CRDs), operators, service mesh

implementations, and production-scale cluster management.

● Advanced proficiency in Linux system administration, performance tuning, and

troubleshooting complex system-level issues.

● Proven experience implementing GitOps workflows using ArgoCD, Flux, or similar tools,

including advanced deployment patterns and progressive delivery.

● Deep understanding of observability principles and hands-on experience with tools such

as Prometheus, Grafana, Datadog, Azure Monitor, or the ELK stack.

● Expert knowledge of networking concepts, including load balancing, CDNs, DNS, VPNs,

service mesh architectures, and distributed systems communication patterns.

● Strong programming and scripting capabilities in Python, Bash, Go, or PowerShell, with

the ability to develop custom tooling and automation frameworks.

● Extensive experience designing and optimizing CI/CD pipelines using Jenkins, GitLab

CI, Azure DevOps, GitHub Actions, or CircleCI.

● Demonstrated ability to lead incident response, conduct root cause analysis, and drive

systemic reliability improvements.

● Excellent communication and leadership skills with proven ability to influence technical

decisions and collaborate with stakeholders at all levels.

● Current certification in AWS (Solutions Architect Associate/Professional or equivalent)

and Azure (Azure Administrator or Azure Solutions Architect), with practical experience

managing production workloads on both platforms.

Good To Have

● Experience with hybrid and multi-cloud networking strategies, including ExpressRoute,

Direct Connect, and cloud interconnects.

● Knowledge of serverless architectures on AWS (Lambda) and Azure (Functions, Logic

Apps) and their operational considerations.

● Proven experience with disaster recovery planning, business continuity, and

implementing multi-region active-active architectures.

● Understanding of machine learning operations (MLOps), data pipeline orchestration, and

supporting ML workloads in production.

● Experience with service mesh technologies such as Istio, Linkerd, or Consul.

● Familiarity with chaos engineering principles and tools like Chaos Monkey or Gremlin.

● Experience with configuration management at scale and policy-as-code tools like Open

Policy Agent (OPA).

● Knowledge of FinOps principles and cloud cost optimization strategies.

Benefits

Why Work at Synthlane?

Real production infrastructure exposure (not just toy tasks). Work closely with engineers building real solutions and client deployments. Learn what “high reliability” actually means in production. Fast-paced, high-impact environment where good work is noticed quickly.