DBS Bank

Analyst, Infrastructure SRE Engineer, Data Platform, Group Technology

DBS Bank Singapore, Singapore, Singapore

Banking · 10,001+ employees

13 h ago
sre Mid (2-5 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, implement, and manage scalable Kubernetes clusters and infrastructure to support a critical data platform. Additionally, you will automate operational tasks, manage CI/CD pipelines, and lead incident response to ensure high system reliability.

What they look for

Kubernetes GitOps Terraform ArgoCD Jenkins Infrastructure as Code Prometheus Grafana OpenTelemetry Docker Python Bash Ansible Helm Kustomize CI/CD

Requirements

Candidates must hold a bachelor's degree in Computer Science or a related field and possess strong experience with Kubernetes and containerization. Proficiency in infrastructure as code tools, scripting, and CI/CD workflows is essential for this role.

Full description

Role Overview: As an Infrastructure Site Reliability Engineer (SRE) specializing in Kubernetes and GitOps, you will be instrumental in ensuring the reliability, availability, and performance of our critical Kubernetes-based data platform. You will apply software engineering principles to operations, focusing on system resilience, automation of operational tasks, and proactive incident prevention. This role requires a strong understanding of infrastructure as code and a commitment to continuous improvement of our production systems.

Key Responsibilities:

  • Design, implement, and manage highly reliable and scalable Kubernetes clusters and underlying infrastructure to support our data platform.
  • Implement and optimize GitOps workflows using ArgoCD for declarative continuous deployment and infrastructure configuration management, ensuring consistency and auditability.
  • Develop and maintain robust CI/CD pipelines using tools like Jenkins, focusing on infrastructure changes and automated validation for reliability and efficiency.
  • Define, monitor, and report on Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to drive continuous improvement in system reliability.
  • Implement comprehensive monitoring, logging, and alerting solutions (e.g., Prometheus, Grafana, OpenTelemetry) to proactively detect and diagnose infrastructure and application issues.
  • Automate infrastructure provisioning, configuration, and management using Infrastructure as Code (IaC) tools such as Terraform, Helm, Kustomize, and Ansible.
  • Lead incident response, root cause analysis (RCA), and post-incident reviews for critical infrastructure outages, implementing preventative measures to enhance system resilience.
  • Collaborate with development, operations, and data engineering teams to embed reliability practices throughout the software development lifecycle.
  • Manage Bitbucket repositories, ensuring version control, code review, and best practices for infrastructure-as-code.
  • Continuously identify and eliminate toil through automation and process improvements.

Job Requirements

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • Strong experience with Kubernetes and containerization technologies (Docker, Podman).
  • Proficiency in Git version control system and experience with Git branching strategies.
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
  • Familiarity with infrastructure as code tools like Terraform, Helm, Kustomize, and Ansible.
  • Solid understanding of networking, security, and monitoring concepts in Kubernetes environments.
  • Strong scripting and automation skills (Bash, Python, or similar).
  • Excellent problem-solving and communication skills.
  • Ability to work effectively in a fast-paced, collaborative environment.

Nice to Have:

  • Certification in Kubernetes (CKA, CKAD) or related technologies.
  • Experience with observability platforms like Prometheus, Grafana, Splunk, or OpenTelemetry.
  • Knowledge of data processing frameworks like Apache Spark, Kafka, or Flink.

Location:

DBS Asia Hub

Job:

Analytics, Technology

Schedule:

Regular

Employee Status:

Full time

Similar roles