N

Senior Site Reliability Engineer

Nordic Investin Group Stockholm, Maine, United States

IT Services and IT Consulting · 201-500 employees

8 h ago
sre Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will define service level objectives, improve telemetry, and automate operational tasks to ensure system reliability. Additionally, you will lead incident responses and collaborate with development teams to influence architecture through production evidence.

What they look for

Site Reliability Engineering Python Go Java Cloud Kubernetes Infrastructure Automation Metrics Logging Tracing Alerting Incident Response Capacity Planning Disaster Recovery Chaos Engineering Resilience Testing

Requirements

Candidates must have strong production engineering or SRE experience and proficiency in programming languages like Python, Go, or Java. Knowledge of cloud infrastructure, Kubernetes, and experience with monitoring and alerting systems is essential.

Full description

On behalf of a partner company, Nordic Investin is looking for a Senior Site Reliability Engineer. You turn reliability from a vague ambition into explicit service targets, useful telemetry and engineering work with clear priorities.

The partner runs digital services where availability and latency directly affect customers. You will work with development teams to improve resilience, incident response and the systems used to understand production behaviour.

The platform is treated as a service used by engineers and business teams. Success will be measured through safer change, lower operational friction and clearer ownership rather than the number of tools introduced.

How you will work

The work combines planned platform development with investigation of real production behaviour. You will collaborate with application teams, security and operations, using their feedback to decide what should become a shared service, an automated control or clear documentation. Ownership continues after the first release.

As a senior colleague, you will own substantial outcomes and help others make stronger decisions. You are expected to recognise risk early, communicate it without drama and move work forward with practical alternatives. The role still includes hands on delivery; seniority here means broader judgement, not distance from the work.

What you will do

  • Define service level indicators and objectives with product teams.
  • Improve telemetry, alert quality and operational dashboards.
  • Automate recurring operational tasks and recovery procedures.
  • Lead or support incident response and blameless reviews.
  • Test capacity, failure modes and disaster recovery.
  • Influence architecture through production evidence.

What you will bring

  • Strong production engineering or SRE experience.
  • Good software development ability in Python, Go, Java or similar.
  • Cloud, Kubernetes and infrastructure automation knowledge.
  • Experience designing metrics, logs, traces and alerts.
  • Calm incident leadership and structured problem solving.
  • Professional English and collaborative communication.

Experience that would add value

  • High traffic consumer or transaction platforms.
  • Chaos engineering and resilience testing.
  • Experience establishing SRE practices in a product organisation.

A background that can succeed here

You may have grown from operations, software engineering or a platform team. What matters is that you can improve reliability through code and design, and that you understand both the technical and human sides of incidents.

What makes the opportunity interesting

The partner offers services with enough scale for reliability work to matter and enough openness for you to change how it is done. Your improvements will be visible in customer experience and engineering focus.

The exact partner, employment model, compensation, start date and working arrangement will be explained openly during the process. Nordic Investin will make sure you understand the context, expectations and decision path before you are asked to commit significant time.

The recruitment conversation

During the process, Nordic Investin will focus on concrete decisions you have made: the context you received, the alternatives you considered, the result you observed and what you would change today. You do not need every optional technology if your core experience transfers and you can explain how you would close the gap.

How to apply

Apply with your CV or LinkedIn profile and a short note describing the most relevant system, product or transformation you have helped deliver. Nordic Investin welcomes candidates with different routes into technology and assesses applicants on relevant capability, judgement and potential.

Similar roles