Senior Site Reliability Engineer
GetFrankly Cluj-Napoca, Romania
Executive Search Services · 2-10 employees
About the role
Define and implement SLIs, SLOs, and error budgets while improving observability and reliability across production services. Participate in incident management, automate operational tasks, and collaborate with engineering squads to foster proactive reliability practices.
What they look for
Requirements
Requires 5-8 years of experience in SRE principles, cloud infrastructure, and production systems at scale. Proficiency in Kubernetes, Terraform, Datadog, and programming languages like Python or Go is essential.
Benefits
Full description
Core details
- Company type: International Retail Group – In-House Product Engineering
- Location: Cluj-Napoca, Romania
- Work model: Hybrid
- Engagement: Employment Contract, Permanent, Full-Time
- Role level: Senior level
- Experience level: 5–8 Years
- Key Skills / Tech Stack: SRE (SLIs, SLOs, error budgets), AWS / GCP / Azure, Kubernetes, Docker, Terraform, GitLab CI/CD, Datadog, Python / Go / Bash
- Department: Engineering
- Industry: Retail – Home Improvement
About the client
Our client is a large international retail group with an in-house engineering organisation, technology built internally, not outsourced. Teams own what they design and run, end to end, on platforms that serve the group's markets and brands across Europe.
The Cluj engineering hub is part of a wider international ecosystem, and the local capability continues to grow.
About the role
This is a hands-on SRE role inside a product engineering environment. You will work directly with software engineering squads that build and own their services in production, helping them improve how those services are designed, observed and operated.
The role sits at the intersection of software engineering, cloud infrastructure and production reliability. You will have hands-on ownership around SLOs, observability, incident response, automation and resilience, while also influencing engineering teams and contributing to shared reliability practices across the organisation. You will collaborate closely with Product Engineering, Platform Engineering, Incident Management, Security, Network and Observability teams.
This is not a role focused on maintaining infrastructure or CI/CD pipelines. The focus is on engineering reliability into digital products used by customers at scale.
Responsibilities
Reliability engineering
- Define and implement SLIs, SLOs and error budgets for production services.
- Help product engineering teams understand and manage reliability trade-offs.
- Identify reliability risks and drive improvements before they become production issues.
- Help design resilient, scalable and highly available cloud-native systems.
- Improve production readiness and operational ownership across engineering squads.
Observability and service health
- Improve observability across services, with a strong focus on customer impact.
- Develop meaningful and actionable monitoring and alerting.
- Use Datadog and related capabilities to improve visibility across distributed systems.
- Help teams move from reactive monitoring towards proactive reliability engineering.
Incident management
- Support the response to major production incidents and participate in the on-call rotation.
- Contribute to effective, blameless post-incident reviews.
- Identify systemic causes rather than treating individual symptoms.
- Ensure post-incident actions translate into long-term engineering improvements.
Automation and operational excellence
- Identify repetitive operational work and reduce engineering toil through automation.
- Improve practices around deployment, monitoring and service operations.
- Contribute to shared SRE standards, tooling and engineering patterns.
- Help make reliability part of everyday software engineering rather than a separate operational activity.
Qualifications
You have strong hands-on experience with production systems and understand that Site Reliability Engineering goes beyond infrastructure automation. Specifically, you bring experience with:
- SRE principles and practices: SLIs, SLOs, error budgets, observability, incident response and automation
- Operating production systems at scale in cloud environments, and the reliability, availability and scalability trade-offs involved in distributed systems
- At least one major cloud platform: AWS, GCP or Azure
- Kubernetes and Docker
- Infrastructure as Code, particularly Terraform
- CI/CD environments such as GitLab CI/CD (preferred), GitHub Actions or Jenkins
- Observability platforms, ideally Datadog
- Scripting or programming in Python, Go, Bash or JavaScript
- Troubleshooting complex production issues
Just as importantly, you are comfortable working directly with software engineers, challenging existing practices where needed, and communicating reliability topics clearly across teams.
Experience defining or owning SLOs for production services, designing alerting around customer impact, taking an active role in major incident management, or introducing SRE practices across multiple engineering teams would be especially valuable. You do not need to tick every technology box ,strong SRE fundamentals and evidence that you have improved the reliability of real production systems matter more than any single tool.
About the offer
Our client offers a competitive package and room to grow:
- Annual performance bonus and employee referral bonus
- Private medical coverage through Regina Maria – Priority Plan, extendable to your spouse or children at no additional cost
- Life insurance through Metropolitan Life
- Eyeglasses vouchers and a co-funded 7Card fitness membership
- Access to LinkedIn Learning, LEO Learning and Bookster
- Meal vouchers, plus gift vouchers on several occasions throughout the year
- 21 days of annual leave, increasing with tenure up to 25 days
- Additional days off when public holidays fall on weekends
- Paid leave for special life events
- Hybrid working model and a flexible schedule
- Modern, collaborative workspace with fresh fruit and premium coffee
You would join an international product engineering organisation where teams build, own and operate their own digital products, and where reliability is treated as an engineering discipline, with real customer impact and the scope to shape SRE practices across a large-scale environment.
Why GetFrankly?
At GetFrankly we are guiding talent and creating futures.
We understand that a career move is not just about a new role , it's about finding a place where your skills, ambitions, and values align.
Give us a call to discuss what's important for your career and future, and we'll try our best to get you involved in interesting projects and provide you with fulfilling career pathways.
We will guide you towards the right environment where your abilities will thrive and have a significant impact.
Similar roles
-
SRE (Site Reliability Engineer) - H/F
Devoteam Levallois-Perret, Ile-de-France, France
-
Ingénieur Observabilité / SRE - H/F
Devoteam Levallois-Perret, Ile-de-France, France
-
SRE
Radware Tel-Aviv, Tel-Aviv District, Israel
-
SRE [Antifraud]
Plata Card Osnabrück, Lower Saxony, Germany
-
Senior SRE & Monitoring Developer
Ford Motor Company India
-
Site Reliability Engineer (SRE) - Early Talent
Nebius Amsterdam, North Holland, Netherlands