Aurora Engineering AB

Senior SRE / Site Reliability Engineer

Aurora Engineering AB Gothenburg, Sweden

IT Services and IT Consulting · 11-50 employees

6 h ago
sre Mid (2-5 yrs) Full-time Sweden
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Senior SRE will design, implement, and maintain highly available production systems while managing incident response and observability. They will also develop automation to improve operational efficiency and collaborate with development teams to enhance system scalability.

What they look for

Site Reliability Engineering Cloud Infrastructure Linux Unix Monitoring Logging Alerting Observability CI/CD Deployment Automation Python Bash Docker Kubernetes Infrastructure Automation Incident Management

Requirements

Candidates must have 4+ years of professional SRE experience, strong cloud infrastructure knowledge, and proficiency in Linux/Unix environments. Additionally, the role requires fluency in Chinese and English, with the candidate currently based in Sweden.

Full description

We are looking for an experienced Senior SRE / Site Reliability Engineer to join our engineering team in Sweden. The ideal candidate will have 4+ years of hands-on SRE experience and a strong background in cloud infrastructure, reliability engineering, automation, monitoring, and incident management.

Key Responsibilities

  • Design, implement, and maintain highly available and reliable production systems.
  • Monitor system health, availability, performance, and capacity.
  • Develop automation to improve operational efficiency and reduce manual activities.
  • Manage incident response, troubleshooting, root-cause analysis, and production issues.
  • Establish and maintain SLIs, SLOs, and SLAs.
  • Build and improve monitoring, logging, alerting, and observability solutions.
  • Support cloud infrastructure, deployments, and production environments.
  • Collaborate with software development and platform teams to improve system reliability.
  • Implement and maintain CI/CD pipelines and deployment automation.
  • Identify performance bottlenecks and proactively improve system scalability and resilience.
  • Contribute to disaster recovery, business continuity, and reliability strategies.
  • Document operational procedures, architecture, incidents, and solutions.
  • Participate in an Agile/DevOps working environment.

Mandatory Requirements

  • 4+ years of professional SRE / Site Reliability Engineering experience.
  • Strong experience with cloud platforms and cloud infrastructure.
  • Hands-on experience with Linux/Unix environments.
  • Strong knowledge of monitoring, logging, alerting, and observability.
  • Experience with CI/CD and deployment automation.
  • Strong scripting/automation skills using technologies such as Python, Bash, or similar.
  • Experience with Docker and Kubernetes.
  • Knowledge of infrastructure automation and configuration management.
  • Strong troubleshooting and root-cause analysis skills.
  • Experience with production incident management and on-call environments.
  • ⭐ Important Requirements:
  • Currently based in Sweden
  • Chinese-speaking
  • Professional working proficiency in English
  • Able to work effectively in an international engineering environment

Similar roles