Q

Senior Site Reliability Engineer

QuTwo Helsinki, Uusimaa, Finland

19 h ago
sre Senior (5-10 yrs) Full-time Finland
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and operate a multi-cloud Kubernetes-based platform while establishing observability, reliability practices, and infrastructure as code. Additionally, you will own the scheduling and cost efficiency of distributed ML workloads and implement security-first infrastructure standards.

What they look for

Kubernetes Site Reliability Engineering Terraform Helm GitOps Python Go Distributed Systems Cloud Infrastructure Observability CI/CD Security Ray Machine Learning Pipelines Infrastructure as Code Autoscaling

Requirements

Candidates must have at least 7 years of experience in SRE or infrastructure roles with deep hands-on expertise in Kubernetes and multi-cloud environments. Proficiency in infrastructure-as-code tools like Terraform and programming languages such as Python or Go is required.

Full description

We are looking for a Senior Site Reliability Engineer to lay the foundations of a dependable, secure, and portable platform across multiple clouds.

Our product combines ML pipelines with small, security-critical SaaS services. We're building toward our first production deployments, so this is a chance to shape how reliability is done here from the start, rather than inherit someone else's choices. One week you might be designing how distributed ML workloads on Ray are scheduled and scaled on Kubernetes. The next you might be setting up observability or defining the first service-level objectives with the team.

We are not looking for someone who has used every tool in our stack. We want an engineer with the systems judgement to decide how our infrastructure should be built and why, and the depth to debug it and take responsibility for how it behaves.

Role Description

  • Building a multi-cloud platform. Design and operate Kubernetes-based infrastructure that runs consistently across multiple clouds, with sensible abstractions where portability matters.
  • Observability and reliability practices. Build the metrics, logging, tracing, and alerting we need, and help define pragmatic SLOs and incident practices that grow with the product.
  • Infrastructure as code and GitOps. Make every environment reproducible, reviewable, and automated end to end.
  • Running ML workloads well. Own scheduling, autoscaling, and cost efficiency for distributed training and inference workloads, including Ray clusters on Kubernetes.
  • Security by design. Own secrets management, network policy, workload identity, and supply-chain security.
  • Enabling the team. Build paved roads for CI/CD and deployment that make the reliable path the easy path for engineers and coding agents alike, and document decisions clearly enough that both can act on them.

Requirements

  • At least 7 years of experience building and operating production systems, including several years in an SRE, platform, or infrastructure role.
  • Deep, hands-on Kubernetes experience, including cluster operations, networking, storage, and troubleshooting.
  • Experience running systems across more than one cloud provider, and a clear view of the trade-offs involved.
  • Strong infrastructure-as-code and automation skills (e.g. Terraform, Helm, GitOps), and fluency in Python or Go for tooling.
  • Solid grasp of distributed systems, their failure modes, and the trade-offs behind reliability targets.
  • Security-first mindset: you think about trust boundaries, identity, secrets, and supply chain by default.
  • Comfort with AI-assisted development: you use coding agents well and review their output critically, including infrastructure changes.
  • Fluent in English, with a proven ability to thrive in agile, cross-functional teams.

Experience in one or more of the following areas is a plus

  • Operating Ray or other distributed compute frameworks in production.
  • Running GPU workloads, including scheduling and utilization optimization.
  • Working closely with ML researchers or running ML pipelines in production.
  • Meeting formal security or compliance requirements (e.g. ISO 27001, SOC 2) in a small company.
  • Cloud cost management and FinOps practices.
  • Familiarity with quantum computing concepts at a high level, or curiosity about them.

Similar roles