About the role
The Site Reliability Engineer will operate and improve production environments while automating infrastructure management across cloud and hybrid platforms. They will collaborate with cross-functional teams to ensure platform resilience, security compliance, and operational excellence.
What they look for
Requirements
Candidates must have proven experience as an SRE, DevOps, or Systems Engineer with strong operational expertise in Microsoft Azure. Proficiency in Infrastructure as Code tools like Terraform, identity management, and security practices is required.
Full description
This is a remote position.
We are looking for a Site Reliability Engineer to join our team and help drive the reliability, scalability, security, and automation of modern cloud platforms.
In this role, you will be responsible for operating and improving production environments, automating infrastructure management, and ensuring platform resilience across cloud and hybrid environments. You will work closely with Engineering, Security, and Platform teams to reduce operational overhead, improve reliability, and accelerate cloud transformation initiatives.
The role combines cloud engineering, infrastructure automation, security, and operational excellence, with a strong focus on Microsoft Azure and Infrastructure as Code practices.
Responsibilities:
- Collaborate with internal and external stakeholders to ensure the successful delivery of infrastructure and platform initiatives
- Deploy, maintain, and scale cloud and hybrid infrastructure environments
- Build, secure, and operate scalable cloud platforms, primarily in Microsoft Azure
- Manage compute, networking, and storage resources across production environments
- Partner with security teams to:
- Identify vulnerabilities
- Implement remediation actions
- Deploy security controls and endpoint protection solutions
- Ensure compliance with security standards
- Administer and optimize identity and access management solutions, including:
- Microsoft Entra ID (Azure AD)
- SSO configurations
- User permissions and access controls
- Develop and maintain Infrastructure as Code (IaC) solutions using Terraform
- Automate provisioning, configuration management, and operational recovery processes
- Implement and maintain monitoring and logging solutions to ensure high availability and rapid incident resolution
- Support FinOps initiatives through cost awareness, resource tagging, and governance practices
- Participate in an on-call rotation to ensure platform reliability and operational continuity
Requirements
- Proven experience as a:
- Site Reliability Engineer (SRE)
- DevOps Engineer
- Systems Engineer supporting cloud production environments
- Strong operational experience with public cloud platforms, particularly:
- Microsoft Azure
- Experience with identity and access management technologies:
- Microsoft Entra ID (Azure AD)
- Single Sign-On (SSO)
- User and permission management
- Experience implementing and managing security tools such as:
- Endpoint Detection & Response (EDR)
- Vulnerability scanners
- Strong experience with:
- Terraform
- Infrastructure as Code (IaC)
- Infrastructure automation
- Experience with:
- Ansible
- Configuration management
- Server provisioning and patching
- Strong troubleshooting and problem-solving skills
- Ability to work autonomously while collaborating effectively with cross-functional teams
- Excellent written and verbal communication skills
- Fluency in English
Nice-to-have
- Experience with container orchestration platforms:
- Kubernetes
- Azure Kubernetes Service (AKS)
- Experience managing CI/CD pipelines using:
- GitLab
- GitLab CI/CD
- Familiarity with Atlassian tools:
- Jira
- Jira Service Management (JSM)
- Confluence
- Experience with observability platforms:
- Prometheus
- Grafana
- Loki
- Azure certifications, such as:
- AZ-104: Azure Administrator Associate
If this sounds like you, share your CV with us and let’s talk!
Similar roles
-
Director, Site Reliability Engineering
Anduril Industries Costa Mesa, California, United States · $253K–$336K/yr
-
Senior Site Reliability Engineer (SRE)
Mirantis Sofia, Sofia-City, Bulgaria
-
Software Engineer III, Site Reliability Engineering
Google Pittsburgh, Pennsylvania, United States · $147K–$210K/yr
-
Senior Software Engineer, Site Reliability Engineering
Google Kirkland, Washington, United States · $174K–$252K/yr
-
Cloud Engineer / Site Reliability Engineer (SRE)
DFDS Denmark Copenhagen, Capital Region of Denmark, Denmark
-
Senior Software Developer, Site Reliability
Google Waterloo, Ontario, Canada · CA$182K–CA$186K/yr