System Reliability Engineer (Application Support + Automation)
Fulcrum Digital · Lisbon, Portugal
IT Services and IT Consulting · 1,001-5,000 employees
About the role
The role involves managing production environments, defining monitoring strategies, and automating processes to ensure system reliability and performance. You will also support CI/CD pipelines, handle incident responses, and collaborate with global teams to optimize service lifecycles.
What they look for
Requirements
Candidates should have experience with Linux shell scripting, ITIL/ITSM frameworks, and CI/CD tools like Jenkins and Ansible. Proficiency in application troubleshooting and monitoring tools such as Splunk or Dynatrace is required.
Full description
Who are we
Fulcrum Digital is an agile and next-generation digital accelerating company providing digital transformation and technology services right from ideation to implementation. These services have applicability across a variety of industries, including banking & financial services, insurance, retail, higher education, food, healthcare, and manufacturing.
The Role
- Plan, manage,
and oversee all aspects of a Production Environment
- Define
strategies for Application Performance Monitoring, Optimization in Prod environment
- Respond to
Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
- Support
deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
- Design, develop and standardize Monitoring and
Alerting mechanism for the supported applications.
- Take a
holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
- Engage in and
improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
- Analyse ITSM
activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
- Support
services before they go live through activities such as system design consulting, capacity planning and launch reviews.
- Support the
application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
- Maintain
services once they are live by measuring and monitoring availability, latency, and overall system health.
- Scale systems
sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
- Work with a
global team spread across tech hubs in multiple geographies and time zones.
- Ability to
share knowledge and explain processes and procedures to others.
- Able to
perform on-call duties on a rotational basis.
- Occasional
off hours work required.
Requirements
- Linux
- Shell
Scripting
- ITIL / ITSM
- SQL - good to have
- Application
Troubleshooting
- Any
Monitoring tool (Preferred Splunk/Dynatrace)
- Jenkins -
CI/CD
- Ansible
- Groovy
Scripting/Yaml
- Git basic/bit
bucket
Good To Have
- Payments Flows, Switching, Settlements, Authorisation flows.
- Even
Framework architecture