Visa

Sr. SRE

Visa Singapore

IT Services and IT Consulting · 10,001+ employees

13 h ago
sre Principal (10+ yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Sr. Site Reliability Engineer manages CI-CD pipelines, infrastructure stability, and server lifecycle maintenance across development and production environments. They also provide hands-on support for cloud infrastructure, containerized services, and incident response to ensure high availability and operational efficiency.

What they look for

Site Reliability Engineering CI-CD AWS GCP Kubernetes Docker Terraform Ansible CloudFormation Python Bash PowerShell Infrastructure as Code Release Management Incident Response System Monitoring

Requirements

Candidates must have at least 8 years of relevant work experience, with a strong background in platform-infrastructure, cloud operations, and Kubernetes. Proficiency in infrastructure-as-code tools, scripting, and production incident management is essential for this role.

Full description

About Us Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.

At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.

Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.

Job Description

The Sr. Site Reliability Engineer is responsible for supporting release management, CI-CD pipeline operations, infrastructure stability, and server lifecycle maintenance across development, staging, and production environments. This role focuses on ensuring reliable, secure, and efficient delivery of applications through robust automation, infrastructure management, and operational best practices.

The engineer works closely with development and platform teams to build, maintain, and optimize CI-CD pipelines, enabling automated build, test, and deployment processes. They contribute to release planning, coordination, and execution, ensuring successful and predictable production releases.

This role provides hands-on support for cloud and on-premise infrastructure (AWS, GCP), including server provisioning, patching, upgrades, and ongoing maintenance to ensure high availability, security, and compliance. The engineer applies infrastructure-as-code practices (Terraform, Ansible, CloudFormation) to standardize and scale environment management.

Additionally, the engineer supports containerized services (Docker, Kubernetes), assists in troubleshooting production issues, and ensures operational readiness through system monitoring and performance optimization. While observability tooling is utilized, the focus remains on operational reliability, release execution, and infrastructure health.

All roles require digital fluency, including the ability to leverage emerging technologies such as Generative AI tools (e.g., Claude, ChatGPT, Microsoft Copilot) to improve productivity and streamline workflows.

Key Responsibilities:

Release Management & CI-CD:

  • Support end-to-end release management, including planning, coordination, and validation of production deployments.
  • Build, maintain, and optimize CI-CD pipelines for automated build, test, and deployment.
  • Troubleshoot and resolve pipeline failures, deployment issues, and release blockers.
  • Partner with development teams to improve release efficiency and reduce deployment risk.

Infrastructure & Server Lifecycle Management:

  • Manage and support cloud infrastructure (AWS, GCP) ensuring availability, reliability, and security.
  • Perform server patching, upgrades, and maintenance in line with security and compliance requirements.
  • Provision and manage infrastructure using infrastructure-as-code tools (Terraform, Ansible, CloudFormation).
  • Ensure environment consistency across development, staging, and production.

Containerized Service Support:

  • Support implementation and operations of containerized services (Docker, Kubernetes).
  • Maintain platform stability and optimize performance of hosted applications.

Operations & Incident Support:

  • Monitor system health and performance, proactively identifying and resolving issues.
  • Participate in incident response, troubleshooting, and root cause analysis for production systems.
  • Provide first-level support for infrastructure and deployment issues, escalating as needed.

Scripting & Automation:

  • Develop and maintain automation scripts to streamline infrastructure operations, deployments, and routine maintenance tasks.
  • Use scripting languages such as Python, Bash, or PowerShell to improve operational efficiency and reduce manual intervention.
  • Automate server patching, environment provisioning, health checks, and deployment workflows.
  • Integrate scripts into CI-CD pipelines to support end-to-end automation.
  • Collaborate with teams to identify automation opportunities and standardize reusable scripts and tooling.

Automation & Continuous Improvement:

  • Identify opportunities to automate repetitive operational tasks and improve workflows.
  • Drive improvements in deployment processes, infrastructure reliability, and operational efficiency.

Documentation & Standards:

  • Create and maintain documentation for release processes, infrastructure configurations, and operational procedures.
  • Ensure adherence to operational and security best practices.

Visa requires at least 3 days in office, expectations of these days will be confirmed by your Hiring Manager.

Qualifications

Education & Experience:

  • Bachelor's degree in a relevant field plus 8+ years of relevant work experience, OR
  • Advanced degree (Master's, MBA, etc.) plus 5+ years of relevant work experience, OR
  • PhD plus 2+ years of relevant work experience, OR
  • 11+ years of relevant work experience without the degree path specified
  • A strong candidate would typically have 8+ years of platform-infrastructure experience, proven ownership of CI-CD and cloud operations, and hands-on expertise with Kubernetes, Terraform, scripting, and production incident management

Visa is an EEO Employer

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability or protected veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with EEOC guidelines and applicable local law.

Similar roles