About the role
The SRE Engineer will contribute to reliability, resilience, and observability objectives by working closely with engineering squads and platform teams. Responsibilities include implementing monitoring solutions, defining SLI/SLOs, and performing root cause analysis to improve service stability.
What they look for
Requirements
Candidates should have extensive knowledge of Linux or Windows, experience with CI/CD standards, and familiarity with monitoring tools like Prometheus and Grafana. Strong analytical skills and experience with both monolithic and distributed application landscapes are required.
Full description
Discover ING Hubs Romania
ING Hubs Romania offers 130 services in software development, data management, non-financial risk & compliance, audit, and retail operations to 24 ING units worldwide, with the help of over 2000 high-performing engineers, risk, and operations professionals.
We started out in 2015 as ING’s software development hub, then steadily expanded our range to include more services and competencies. Now we provide borderless services with bank-wide capabilities and operate from Bucharest.
Our tech capabilities remain the core of our business, with more than 1800 colleagues active in Data and Analytics Tech, Tech Foundation and Channels, Retail Core Banking and Architecture, and Global Products and Technology Services.
We enjoy a flexible way of working and a highly collaborative environment, where fair and constructive feedback is encouraged.
For us, impact isn't a perk. It's the driver of our work. We are guided and rewarded by a shared desire to make the world a better place, one innovative solution at a time. Our colleagues make it their job to do impactful things and they love doing it in good company. Do you?
Here’s a sneak peak of what our colleagues say about working within ING Hubs Romania:
- At ING, we're building the solutions of tomorrow, today | 80% of our colleagues in Romania agree
The Mission
ING’s Payment and Settlement Services (“PSS”) aims to further mature, develop and expand ING’s state of the art payments platforms and settlements services. The PSS focuses on delivering payments and settlement services, meeting the expectations of ING’s business lines and beyond, whilst ensuring the basics: safe, secure, compliant, and reliable. The PSS is an independent unit within ING’s COO domain, reporting directly into ING’s Chief Operating Officer.
The PSS is responsible for providing standardized payments and settlements services to multiple ING business lines (WB, Retail, and others). There is a clear ambition to also look for external commercialization of our payment platforms.
The PSS changes the organization towards a true Services-based organization and to deliver on ambitious promises, we have organized ourselves into a strong strategic and commercially focused organization around our eight services, where our squads working on those products are enabled by clear prioritization, accountability, and capabilities. And right there, the SRE Engineer comes in.
As an SRE Engineer, you will contribute to achieving the Site Reliability Engineering objectives within PSS. You will work closely with engineering squads, architects, platform teams, and operational stakeholders to improve reliability, resilience, observability, and operational excellence across our services. This role requires strong technical expertise, collaboration skills, and a proactive mindset focused on continuous improvement.
The SRE team supports the adoption and implementation of Site Reliability Engineering practices across the organization. As an SRE Engineer, you will help teams build and operate reliable, scalable, secure, and observable services while contributing to the evolution of SRE capabilities and ways of working within PSS.
Your day to day
As an SRE Engineer you are dedicated to supporting the PSS organization in achieving its reliability engineering objectives. The SRE focus is on creating reliable and available services for customers. This role holds multiple facets and, dependent on the match with your skills, can entail the following activities:
- Support or perform reviews and analysis on Asset/Service implementations regarding their monitoring, observability, and alerting setup.
- Support or perform reviews and analysis on Asset/Service implementations regarding their resilience architecture and operational readiness.
- Implement, maintain, and improve monitoring, alerting, logging, and observability solutions using ING standards and industry best practices.
- Contribute to the definition, implementation, and monitoring of SLI/SLOs, error budgets, and availability reporting.
- Analyze and support Root Cause Analysis and Post-Mortems, helping identify and implement improvement actions to prevent future incidents.
- Support engineering teams in troubleshooting complex production issues and improving service stability.
- By analyzing Incident, Problem, and Change data, identify improvement opportunities and contribute to the implementation of structural solutions.
- Identify operational toil and contribute to automation initiatives that improve efficiency, reliability, and operational excellence.
- Support resilience testing, disaster recovery exercises, and operational readiness assessments.
- Provide hands-on support when application teams require guidance implementing monitoring, observability, alerting, or reliability standards.
- Collaborate closely with development teams throughout the software lifecycle to improve reliability, availability, scalability, and maintainability.
- Contribute to the continuous improvement of operational processes, runbooks, and engineering practices.
- Share knowledge and best practices through documentation, workshops, and collaboration with colleagues across the organization.
- Contribute to the E&R organisation.
- Participate in Global SRE Guilds and Communities of Practice.
What you'll bring to the team
PSS SRE is looking for a broad skill set for an SRE Engineer, comprising both modern-day and legacy IT stacks. Knowledge of coding, testing, automation, observability, and CI/CD standards, combined with a healthy dose of curiosity, is a given in you. You are not afraid of a challenge no matter what the subject is. Impediments are for you to solve; neglecting them is not an option.
Currently we are looking for someone who feels they fit the above description and has an understanding of the below technical skill set:
- Extensive knowledge of Linux (RHEL) or Windows.
- Good knowledge of SQL and familiarity with RDBMS databases (Oracle, MS SQL) or NoSQL databases (Cassandra).
- Good knowledge of CI/CD standards (preferably Azure DevOps).
- Extensive knowledge of IT tools for sharing and collaborating.
- Experience with monolithic and distributed application landscapes.
- Good knowledge in at least one programming language.
- Good knowledge in Networking and IPv4 and IPv6 stacks.
- Experience with monitoring, observability, and alerting tools (Prometheus, ELK, Grafana, OpenTelemetry, etc.).
- Familiarity with Site Reliability Engineering (SRE) concepts such as SLI, SLO, error budgeting, reliability engineering, and availability reporting.
- Good knowledge of containers and cloud tooling.
- Knowledge and experience in using OpenTelemetry is a bonus.
- Good knowledge of IT security principles.
- Ability to understand customer and engineering needs and translate them into practical solutions.
- Strong analytical and problem-solving mindset.
- Ability to collaborate effectively within cross-functional teams.
- Continuous improvement mindset and willingness to learn new technologies and practices.
- Foreign languages: English (advanced).
If you want to deep dive into the processing of personal data conducted by ING Hubs Romania during the recruitment process and your rights related to it, read the privacy notices on our website (make sure to scroll until you reach the Data Protection section/ Candidates tab).
Similar roles
-
Senior Site Reliability Engineer (AWS / EKS)
Salve.Inno Consulting Philippines
-
Staff Site Reliability Engineer
Coupang Bengaluru, Karnataka, India
-
Software Engineering, Site Reliability Engineering BS/MS Intern, 2027
Google Paris, Ile-de-France, France · €71K/yr
-
Site Reliability Engineer
Experian Sepang, Selangor, Malaysia
-
Site Reliability Engineer
Delinea Bracknell, England, United Kingdom · £70K–£85K/yr
-
Senior Resident Engineer (SRE) - Mechanical & Electrical #SGINFRA
Bureau Veritas Singapore, Singapore