Claranet

Senior Site Reliability Engineer

Claranet Lisbon, Portugal

Information Technology & Services · 1,001-5,000 employees

19 h ago
sre Mid (2-5 yrs) Full-time Portugal
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves ensuring the reliability, performance, and scalability of cloud-based platforms through operational support and incident management. Additionally, it requires designing and implementing automation, observability, and infrastructure improvements to enhance service resilience.

What they look for

Azure Kubernetes Terraform CI/CD Datadog Python Bash PowerShell Infrastructure as Code Observability Site Reliability Engineering Automation Cloud Engineering Microservices Troubleshooting Communication

Requirements

Candidates must have at least 4 years of experience in SRE, DevOps, or Platform Engineering with hands-on expertise in Azure and Kubernetes. A degree in Computer Science or a related field is required, along with proficiency in Infrastructure as Code tools and scripting languages.

Benefits

Regular professional development Certification paths resources Regular teambuilding programs Friendly workplace

Full description

We're fast learners, hard workers, natural collaborators... and we Make Modern Happen!

Our ambition is to unlock the potential of our digital world so that organisations everywhere can innovate and thrive securely.

We aim to achieve this goal by bringing together the world’s most talented people and the most powerful technologies, combining them to address our customers' challenges and to build something stronger together.

If you share our vision, join us!

We are looking for a Site Reliability Engineer to join our team and help us build and operate reliable, scalable and secure technology platforms.

This role combines two complementary areas of work:

  • 50% Operational Excellence: ensuring the smooth

operation of our platforms, responding to service requests and incidents, troubleshooting issues and continuously improving reliability.

  • 50% Engineering & Improvement

Projects: designing and implementing automation, observability, infrastructure and platform improvements that make our services more resilient and easier to operate.

This role is responsible for ensuring the reliability, performance, security, and scalability of cloud-based platforms, primarily in Azure and Kubernetes environments. The position combines operational support, infrastructure engineering, automation, and Site Reliability Engineering (SRE) practices.

Your responsibilities include:

  • Monitoring and

maintaining cloud and Kubernetes platforms to ensure high availability and performance.

  • Investigating and

resolving incidents, conducting root cause analysis, and driving continuous service improvements.

  • Designing, deploying,

and managing scalable infrastructure in Azure.

  • Managing Kubernetes

environments, preferably with AKS (Azure Kubernetes Service).

  • Developing and

maintaining Infrastructure as Code using Terraform.

  • Building and improving

CI/CD pipelines and automating operational processes.

  • Implementing

observability solutions, including monitoring, logging, tracing, and alerting tools such as Datadog.

  • Defining and tracking

reliability and performance metrics (SLIs, SLOs, and error budgets).

  • Collaborating with

development and infrastructure teams to deliver reliable, secure, and maintainable platform solutions.

  • Promoting DevOps,

automation, knowledge sharing, and a culture of continuous improvement.

You must have:

  • Degree in Computer Science, Engineering or a related field, or

equivalent practical experience.

  • At least 4 years of experience in SRE, DevOps, Platform

Engineering, Cloud Engineering or a similar role.

  • Hands-on experience with Microsoft Azure, particularly compute,

networking and storage services.

  • Practical experience with Kubernetes; experience with AKS is an

advantage.

  • Experience with Terraform or another Infrastructure as Code tool.
  • Familiarity with CI/CD practices and version control systems.
  • Experience with monitoring, logging and alerting platforms such as

Datadog, Azure Monitor, Prometheus, Grafana or equivalent.

  • Good scripting skills in Bash, Python or PowerShell.
  • Understanding of software development and deployment practices.
  • Experience with .NET and/or Java, microservices or business

applications deployed on Kubernetes is a strong advantage.

  • Ability to troubleshoot complex technical issues in a structured

and collaborative way.

  • Good written and verbal communication skills in English.

We value:

  • Experience with AWS or Google Cloud.
  • Experience with .NET or Java application development.
  • Knowledge of SLI/SLO frameworks, error budgets and incident

management practices.

  • Experience with distributed systems, APIs and cloud-native

architectures.

  • Familiarity with security, networking and identity concepts in

Azure and Kubernetes.

  • Relevant certifications, such as:
  • Microsoft

Certified: Azure Fundamentals

  • Microsoft

Certified: Azure Solutions Architect Expert

  • Certified

Kubernetes Administrator

  • HashiCorp

Terraform Associate

  • Datadog

Fundamentals

We offer:

  • Regular

professional development;

  • Certification

paths resources;

  • Regular teambuilding programs;
  • Friendly workplace.

Workplace: Lisbon (Hybrid)

Claranet: Make Modern Happen!

Similar roles