Senior Site Reliability Engineer
Claranet Lisbon, Portugal
Information Technology & Services · 1,001-5,000 employees
About the role
The role involves ensuring the reliability, performance, and scalability of cloud-based platforms through operational support and incident management. Additionally, it requires designing and implementing automation, observability, and infrastructure improvements to enhance service resilience.
What they look for
Requirements
Candidates must have at least 4 years of experience in SRE, DevOps, or Platform Engineering with hands-on expertise in Azure and Kubernetes. A degree in Computer Science or a related field is required, along with proficiency in Infrastructure as Code tools and scripting languages.
Benefits
Full description
We're fast learners, hard workers, natural collaborators... and we Make Modern Happen!
Our ambition is to unlock the potential of our digital world so that organisations everywhere can innovate and thrive securely.
We aim to achieve this goal by bringing together the world’s most talented people and the most powerful technologies, combining them to address our customers' challenges and to build something stronger together.
If you share our vision, join us!
We are looking for a Site Reliability Engineer to join our team and help us build and operate reliable, scalable and secure technology platforms.
This role combines two complementary areas of work:
- 50% Operational Excellence: ensuring the smooth
operation of our platforms, responding to service requests and incidents, troubleshooting issues and continuously improving reliability.
- 50% Engineering & Improvement
Projects: designing and implementing automation, observability, infrastructure and platform improvements that make our services more resilient and easier to operate.
This role is responsible for ensuring the reliability, performance, security, and scalability of cloud-based platforms, primarily in Azure and Kubernetes environments. The position combines operational support, infrastructure engineering, automation, and Site Reliability Engineering (SRE) practices.
Your responsibilities include:
- Monitoring and
maintaining cloud and Kubernetes platforms to ensure high availability and performance.
- Investigating and
resolving incidents, conducting root cause analysis, and driving continuous service improvements.
- Designing, deploying,
and managing scalable infrastructure in Azure.
- Managing Kubernetes
environments, preferably with AKS (Azure Kubernetes Service).
- Developing and
maintaining Infrastructure as Code using Terraform.
- Building and improving
CI/CD pipelines and automating operational processes.
- Implementing
observability solutions, including monitoring, logging, tracing, and alerting tools such as Datadog.
- Defining and tracking
reliability and performance metrics (SLIs, SLOs, and error budgets).
- Collaborating with
development and infrastructure teams to deliver reliable, secure, and maintainable platform solutions.
- Promoting DevOps,
automation, knowledge sharing, and a culture of continuous improvement.
You must have:
- Degree in Computer Science, Engineering or a related field, or
equivalent practical experience.
- At least 4 years of experience in SRE, DevOps, Platform
Engineering, Cloud Engineering or a similar role.
- Hands-on experience with Microsoft Azure, particularly compute,
networking and storage services.
- Practical experience with Kubernetes; experience with AKS is an
advantage.
- Experience with Terraform or another Infrastructure as Code tool.
- Familiarity with CI/CD practices and version control systems.
- Experience with monitoring, logging and alerting platforms such as
Datadog, Azure Monitor, Prometheus, Grafana or equivalent.
- Good scripting skills in Bash, Python or PowerShell.
- Understanding of software development and deployment practices.
- Experience with .NET and/or Java, microservices or business
applications deployed on Kubernetes is a strong advantage.
- Ability to troubleshoot complex technical issues in a structured
and collaborative way.
- Good written and verbal communication skills in English.
We value:
- Experience with AWS or Google Cloud.
- Experience with .NET or Java application development.
- Knowledge of SLI/SLO frameworks, error budgets and incident
management practices.
- Experience with distributed systems, APIs and cloud-native
architectures.
- Familiarity with security, networking and identity concepts in
Azure and Kubernetes.
- Relevant certifications, such as:
- Microsoft
Certified: Azure Fundamentals
- Microsoft
Certified: Azure Solutions Architect Expert
- Certified
Kubernetes Administrator
- HashiCorp
Terraform Associate
- Datadog
Fundamentals
We offer:
- Regular
professional development;
- Certification
paths resources;
- Regular teambuilding programs;
- Friendly workplace.
Workplace: Lisbon (Hybrid)
Claranet: Make Modern Happen!
Similar roles
-
Senior Software Engineering Manager, Site Reliability Engineering
Google Sunnyvale, California, United States · $262K–$364K/yr
-
Director, Site Reliability Engineering
Omnicell United Kingdom
-
Associate Site Reliability and Forward Deployed Engineer
Abacus Insights Pune, Maharashtra, India
-
Staff Software Engineer – SRE & AIOps
ServiceNow Vancouver, British Columbia, Canada · CA$126K–CA$220K/yr
-
Site Reliability Engineering (SRE) Manager
McAfee Frisco, Texas, United States
-
Software Engineering Manager, SRE, AI Foundations, Data Intelligence
Google Sydney, New South Wales, Australia