Site Reliability Engineer
Armor Defense Inc Pune, Maharashtra, India
Computer and Network Security · 51-200 employees
About the role
The Site Reliability Engineer will design and maintain resilient hybrid infrastructure, focusing on automation, system reliability, and proactive problem prevention. They will also lead incident response, manage DR strategies, and mentor team members while ensuring compliance with security standards.
What they look for
Requirements
Candidates must have 6-10 years of experience in SRE roles within Windows-based production environments, with deep expertise in VMware and hybrid identity management. Proficiency in scripting languages like Python and PowerShell, along with experience in IaC tools like Terraform and CI/CD pipelines, is required.
Full description
At Armor, we are committed to making a meaningful difference in securing cyberspace. Our vision is to be the trusted protector and de facto standard that cloud-centric customers entrust with their risk. We strive to continuously evolve to be the best partner of choice, breaking norms and tirelessly innovating to stay ahead of evolving cyber threats and reshaping how we deliver customer outcomes. We are passionate about making a positive impact in the world, and we’re looking for a highly skilled and experienced talent to join our dynamic team.
Armor has unique offerings to the market so customers can a) understand their risk b) leverage Armor to co-manage their risk or c) completely outsource their risk to Armor.
Learn more at: https://www.armor.com
SUMMARY
We are looking for a highly skilled Senior Site Reliability Engineer (SRE) to join our infrastructure team with expertise across Cloud Deployments, Microsoft Entra ID (Azure AD), Active Directory, Office 365, Zerto, Rubrik, VMware, and NSX-T. This hands-on, automation-heavy role focuses on system reliability, scalability, and proactive problem prevention. You will be responsible for building and maintaining resilient infrastructure, automating repetitive tasks, monitoring, and improving
performance, and driving incident reduction strategies across hybrid cloud environments.
This role operates in a hybrid structure with on-site presence four days a week, specifically Monday, Tuesday, Wednesday, and Thursday, based in Pune, India.
ESSENTIAL DUTIES AND RESPONSIBILITIES (Additional duties may be assigned as required)
- Own the design and roadmap for hybrid identity (Entra ID and on-prem AD), including conditional access and secure authentication standards.
- Design, deploy, and administer VMware vSphere and NSX-T environments, including capacity planning and secure network topologies.
- Define and own the DR strategy for Zerto and Rubrik, ensuring RTO/RPO and compliance requirements are met.
- Plan and lead workload migrations across VMware, AWS, and Azure.
- Define SLIs/SLOs and design monitoring, alerting, and incident response workflows for production services.
- Build and maintain reusable Terraform modules; set IaC standards for the team.
- Develop automation frameworks in Python/PowerShell, including self-healing and proactive detection.
- Design and maintain CI/CD and GitOps pipelines (GitHub Actions, Azure DevOps or equivalent).
- Own the automated vulnerability patch management program.
- Optimize Office 365 services (Exchange Online, SharePoint, Teams, Intune) for security and performance.
- Lead response and resolution for high-impact incidents; act as on-call escalation point.
- Conduct blameless postmortems and drive remediation to closure.
- Mentor SRE team members and review scripts, Terraform, and change plans.
- Represent infrastructure in security and compliance audits (PCI, HIPAA, ISO).
REQUIRED SKILLS
- 6–10 years in SRE on Windows-based production environments, including ownership of platform reliability.
- Deep expertise in VMware (vSphere, ESXi, vCenter, NSX-T), including capacity planning and architecture design.
- Expert-level Microsoft Entra ID and on-prem Active Directory: hybrid identity design, conditional access strategy, identity protection, and troubleshooting at scale.
- Strong programming/scripting in Python and PowerShell, with the ability to build reusable automation frameworks and self-healing mechanisms. Infrastructure-as-code with Terraform, including module design and DRY principles.
- CI/CD and GitOps experience (GitHub Actions, Azure DevOps or equivalent).
- Designing and implementing observability: defining SLIs/SLOs and building dashboards and alerting with Prometheus, Grafana, Datadog, or ELK.
- DR design and ownership with Zerto and Rubrik, including RTO/RPO planning, DR testing, and compliance.
- Strong networking knowledge: virtual network topology design, firewalls, load balancing, and DNS.
PREFERRED SKILLS
- Production experience with AWS or Azure, including Secure Landing Zone design.
- Experience planning and executing workload migrations across VMware, AWS, and Azure.
- Kubernetes and Linux/Unix production experience.
- Ansible or similar configuration management tools.
- Experience owning vulnerability and patch management programs at scale.
- Experience supporting security and compliance audits (PCI, HIPAA, ISO).
- Relevant certifications (VCP, AZ-104/AZ-305, SC-300, HashiCorp Terraform Associate, or equivalent).
- Prior experience mentoring engineers or leading small technical initiatives.
- Strong troubleshooting skills in complex hybrid environments.
- Strong written and verbal communication; comfortable presenting designs to stakeholders.
WHY ARMOR
Join Armor if you want to be part of a company that is redefining cybersecurity. Here, you will have the opportunity to shape the future, disrupt the status quo, and be a part of a team that celebrates energy, passion, and fresh thinking. We are not looking for someone who simply fills a role – we want talent who will help us write the next chapter of our growth story.
ARMOR CORE VALUES
- Commitment to Growth: A growth mindset that encourages continuous learning and improvement with adaptability in the face of challenges.
- Integrity Always: Sustain trust through transparency + honesty in all actions and interactions regardless of circumstances.
- Empathy In Action: Active understanding, compassion and support to the needs of others through genuine connection.
- Immediate Impact: Taking initiative with swift, informed actions to deliver positive outcomes.
- Follow-Through: Dedication to delivering finished results with attention to quality and detail to achieve the desired outcomes.
WORK ENVIRONMENT
The work environment characteristics described here are representative of those an employee encounters while performing the essential functions of this job. The noise level in the work environment is usually low to moderate. This role is based in our Pune, India office in a hybrid arrangement, with on-site presence required four days a week (Monday through Thursday).
Equal opportunity employer - it is the policy of the company to comply with all employment laws and to afford equal employment opportunity to individuals in all aspects of employment, including in selection for job opportunities, without regard to race, color, religion, sex, national origin, age, disability, genetic information, veteran status, or any other consideration protected by federal, state or local laws.
Similar roles
-
Senior Site Reliability Engineer
Formation Bio San Francisco, California, United States · $186K–$232K/yr
-
Staff Site Reliability Engineer
Okta Dublin, Leinster, Ireland · €92K–€126K/yr
-
Senior Site Reliability Engineer
Okta Dublin, Leinster, Ireland · €76K–€104K/yr
-
Site Reliability Engineer Engineer
Modus Create United States
-
SRE
Zensar Bangalore South, Karnataka, India
-
Staff Site Reliability Engineer
Renesas Electronics San Diego, California, United States · $170K–$210K/yr