SRE (Application Support + Dev-Ops + Automation)
Fulcrum Digital Shankill, Leinster, Ireland
IT Services and IT Consulting · 1,001-5,000 employees
About the role
The Site Reliability Engineer will focus on application support, DevOps, and automation to ensure system reliability and performance. They will collaborate with development teams to resolve technical issues and drive continuous improvement across operational processes.
What they look for
Requirements
Candidates must have strong experience in scripting, cloud infrastructure, and systems administration. Proficiency in DevOps practices, observability tools, and troubleshooting complex technical issues is essential for this role.
Full description
Who We Are
Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.
The Role
We're looking for a Site Reliability Engineer to join our Business Operations team, focusing on application support, DevOps, and automation. In this role, you'll apply your expertise in reliability engineering to keep critical systems running smoothly, resolve issues efficiently, and drive continuous improvement across operational processes. You'll work as an experienced individual contributor, collaborating closely with development teams and stakeholders to ensure that reliability solutions align with both technical requirements and business goals, while also having the opportunity to contribute to new product and service initiatives.
What You'll Do
- Independently execute key elements of projects and processes within the Site Reliability Engineering area, applying in-depth discipline knowledge and best practices to resolve problems and roadblocks
- Evaluate operational requirements and help develop technical solutions within existing frameworks
- Support automation and scripting efforts to improve operational workflows and incident response processes
- Troubleshoot and resolve routine and moderately complex system issues, escalating when necessary to maintain system health
- Contribute to documentation, knowledge sharing, and best practices to enhance team operational procedures
- Collaborate with development teams and stakeholders to ensure reliability solutions align with technical and business needs
- Participate in reviews and quality assurance activities to uphold system stability standards
- Contribute to solution development for new products/services and manage smaller projects or initiatives as an experienced individual contributor
Requirements
Requirements
- Observability skills, using scripting and tooling to collect, analyze, and visualize metrics, logs, and traces for incident detection, diagnosis, and continuous improvement
- Programming and scripting ability (e.g., Python, Go, Bash, or similar) to automate tasks, build operational tools, and support monitoring, deployment, and incident response
- Systems and network administration skills, including configuring, operating, and troubleshooting Linux/Unix systems and network components, with knowledge of networking concepts, protocols, and security
- Cloud computing and infrastructure experience, designing, deploying, and managing applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP) for scalability, security, and operational efficiency
- Understanding of reliability and scalability principles, designing and operating systems for high availability, fault tolerance, and disaster recovery
- Knowledge of DevOps practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery
- Strong troubleshooting capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks
- Capacity planning and performance optimization skills to monitor resource utilization, forecast capacity needs, and optimize performance
- IT service management knowledge, applying principles to incident, problem, and change management for reliable service delivery
- Proactive monitoring mindset, using application reliability signals to anticipate issues and drive preventative improvements
Similar roles
-
Site Reliability Engineer
Armor Defense Inc Pune, Maharashtra, India
-
Senior Site Reliability Engineer
Formation Bio San Francisco, California, United States · $186K–$232K/yr
-
Staff Site Reliability Engineer
Okta Dublin, Leinster, Ireland · €92K–€126K/yr
-
Senior Site Reliability Engineer
Okta Dublin, Leinster, Ireland · €76K–€104K/yr
-
Site Reliability Engineer Engineer
Modus Create United States
-
SRE
Zensar Bangalore South, Karnataka, India