Fulcrum Digital

Sr System Reliability Engineer (Application Support + Automation)

Fulcrum Digital Shankill, Leinster, Ireland

IT Services and IT Consulting · 1,001-5,000 employees

Sep 01
Senior (5-10 yrs) Full-time Ireland
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role involves managing production environments, defining monitoring strategies, and automating processes to ensure system reliability and performance. You will also engage in the full service lifecycle, from design and capacity planning to incident response and mentoring junior team members.

What they look for

Linux Mainframe Shell Scripting ITIL ITSM Application Troubleshooting SQL Splunk Dynatrace Jenkins CI/CD Groovy Yaml Git Bitbucket Ansible

Requirements

Candidates must have strong experience with Linux, Mainframe, Shell Scripting, and ITIL/ITSM frameworks. Proficiency in SQL, CI/CD pipelines, and monitoring tools like Splunk or Dynatrace is required.

Full description

Who are we Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.

The Role

  • Plan, manage,

and oversee all aspects of a Production Environment

  • Define

strategies for Application Performance Monitoring, Optimization in Prod environment

  • Respond to

Incidents and improvise platform based on feedback and measure the reduction of incidents over time.

  • Support

deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.

  • Design, develop and standardize Monitoring and

Alerting mechanism for the supported applications.

  • Take a

holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.

  • Engage in and

improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.

  • Analyze ITSM

activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.

  • Support

services before they go live through activities such as system design consulting, capacity planning and launch reviews.

  • Support the

application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.

  • Maintain

services once they are live by measuring and monitoring availability, latency and overall system health.

  • Scale systems

sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.

  • Work with a

global team spread across tech hubs in multiple geographies and time zones.

  • Ability to

share knowledge and explain processes and procedures to others.

  • Share

knowledge and mentor junior resources

  • Able to

perform on-call duties on a rotational basis.

  • Occasional

off hours work required.

Requirements

Skills –

Must Have

  • Linux
  • Mainframe​
  • Shell

Scripting

  • ITIL / ITSM, Application

Troubleshooting

  • SQL
  • Any

Monitoring tool (Preferred Splunk/Dynatrace)

  • Jenkins -

CI/CD

  • Groovy

Scripting/Yaml - basic

  • Git basic/bit

bucket - basic

  • Ansible/Chef

- good to have

Good To Have

  • Even Framework architecture