Lead Site Reliability Engineer
Swift Transportation Kuala Lumpur, Kuala Lumpur, Malaysia
Transportation/Trucking/Railroad · 10,001+ employees
About the role
The Lead Site Reliability Engineer will oversee system administration, architecture design, and automation to ensure high system reliability and operational excellence. They will also provide technical leadership to a team of up to 10 engineers while collaborating with cross-functional teams to resolve production issues.
What they look for
Requirements
Candidates must have at least 10 years of system administration experience and 2 years of leadership experience. Proficiency in Linux, scripting, automation tools, and distributed systems architecture is required.
Benefits
Full description
ABOUT US
We’re the world’s leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we’re proud to support the global economy.
We’re unique too. We were established to find a better way for the global financial community to move value – a reliable, safe and secure approach that the community can trust, completely. We’re always striving to be better and are constantly evolving in an ever-changing landscape, without undermining that trust. Five decades on, our vibrant community reflects the complexity and diversity of the financial ecosystem. We innovate diligently, test exhaustively, then implement fast. In a connected and exciting era, our mission has never been more relevant. Swift now has a presence in 200+ countries and legal territories to serve a community of more than 12,000 banks and financial institutions.
What to expect:
- Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.
- Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.
- Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence.
- Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.
- Work collaboratively with other team members, provide technical leadership and guidance to a team of up to 10 SRE engineers, driving engineering excellence, reliability, and operational best practices.
- Organize an efficient handover through high quality documentation and training.
- Automate the deployment and operation of multi-tenant infrastructure, handling tasks that ensure system resilience and availability.
- Develop and maintain monitoring tools, dashboards, and self-healing mechanisms.
- Participate in on-call rotations, weekend deployment duty. conduct blameless postmortems, and drive continuous learning.
- Work closely with developers, product teams, and engineering stakeholders to troubleshoot issues, improve systems, and integrate reliability improvements
- Capable of providing accurate project estimates and strategically adapting plans throughout the project lifecycle.
What will make you successful?
- Minimum 10 years of system administration experience in an (preferably) international setting.
- Minimum 2 years of experience leading project or team.
- Familiarity or experience with data ingestion with big data technologies (Elastic Search, Logstash, Kibana and kafka).
- Experience with CICD development & deployment tools such as Maven, Jenkins, Nexus, Git, and Docker.
- Proficiency in Linux OS
- Proficiency in scripting and automation (e.g. Python, PowerShell, YAML) with the ability to develop tools and infrastructure as code (Preferably Ansible, Terraform, Kubernetes, OpenShift).
- Understanding of distributed systems and microservices architectures, including REST and SOAP APIs.
- Hands-on experience with ITIL processes, including Incident, Problem, and Continual Improvement, is an advantage.
- Experience working within an Agile-driven environment.
- Practical experience in building metrics for data-driven reporting.
- Strong interpersonal skills with a customer-centric mindset and ability to work effectively across diverse cultures.
- Proven ability to collaborate with both local and remote teams across different time zones.
- Familiarity with or experience in managing VM hosts using vCenter is an advantage.
- Strong technical leadership to a team of up to 10 SRE engineers, able to balance hands-on expertise with effective mentorship, design reviews, and operational leadership.
What we offer
We give you the freedom to be yourself. We are creating an environment of unique individuals – like you – with different perspectives on the financial industry and the world. A diverse and inclusive environment in which everyone’s voice counts and where you can reach your full potential.
We are committed to an inclusive and accessible recruitment process. If you require a reasonable accommodation related to accessibility during your application or interview, please contact accessibility-Sysgroup@swift.com or indicate this in your application.
Please note that this mailbox is not monitored for general recruitment enquiries and should only be used for accessibility or accommodation-related requests (for example related to vision, hearing or neurodiversity).
All requests are confidential and will not affect your candidacy.
Don’t meet every single requirement? At Swift, we are dedicated to building a workplace where people can bring their full selves and ideas to the team, so if you are excited about this role, we encourage you to apply even if you do not meet every single qualification.
Similar roles
-
Senior SRE Software Engineer - Software Developer Platform
Apple Shanghai, Shanghai, China
-
Senior SRE Software Engineer - ASE Compute
Apple Shanghai, Shanghai, China
-
Senior SRE Engineer - ASE Traffic & Secure Services Network
Apple Shanghai, Shanghai, China
-
Senior SRE Engineer - ASE ADP Compute
Apple Shanghai, Shanghai, China
-
Lead Engineer - SRE
kiwibankpeople Auckland, Auckland, New Zealand
-
Lead Site Reliability Engineer, Chief Digital Office
UnitedHealth Group Eden Prairie, Minnesota, United States · $113K–$193K/yr