Analyst, Infrastructure SRE Engineer, Data Platform, Group Technology
DBS Bank Singapore, Singapore, Singapore
Banking · 10,001+ employees
About the role
You will design, implement, and manage scalable Kubernetes clusters and infrastructure to support a critical data platform. Additionally, you will automate operational tasks, manage CI/CD pipelines, and lead incident response to ensure high system reliability.
What they look for
Requirements
Candidates must hold a bachelor's degree in Computer Science or a related field and possess strong experience with Kubernetes and containerization. Proficiency in infrastructure as code tools, scripting, and CI/CD workflows is essential for this role.
Full description
Role Overview: As an Infrastructure Site Reliability Engineer (SRE) specializing in Kubernetes and GitOps, you will be instrumental in ensuring the reliability, availability, and performance of our critical Kubernetes-based data platform. You will apply software engineering principles to operations, focusing on system resilience, automation of operational tasks, and proactive incident prevention. This role requires a strong understanding of infrastructure as code and a commitment to continuous improvement of our production systems.
Key Responsibilities:
- Design, implement, and manage highly reliable and scalable Kubernetes clusters and underlying infrastructure to support our data platform.
- Implement and optimize GitOps workflows using ArgoCD for declarative continuous deployment and infrastructure configuration management, ensuring consistency and auditability.
- Develop and maintain robust CI/CD pipelines using tools like Jenkins, focusing on infrastructure changes and automated validation for reliability and efficiency.
- Define, monitor, and report on Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to drive continuous improvement in system reliability.
- Implement comprehensive monitoring, logging, and alerting solutions (e.g., Prometheus, Grafana, OpenTelemetry) to proactively detect and diagnose infrastructure and application issues.
- Automate infrastructure provisioning, configuration, and management using Infrastructure as Code (IaC) tools such as Terraform, Helm, Kustomize, and Ansible.
- Lead incident response, root cause analysis (RCA), and post-incident reviews for critical infrastructure outages, implementing preventative measures to enhance system resilience.
- Collaborate with development, operations, and data engineering teams to embed reliability practices throughout the software development lifecycle.
- Manage Bitbucket repositories, ensuring version control, code review, and best practices for infrastructure-as-code.
- Continuously identify and eliminate toil through automation and process improvements.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, or related field.
- Strong experience with Kubernetes and containerization technologies (Docker, Podman).
- Proficiency in Git version control system and experience with Git branching strategies.
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Familiarity with infrastructure as code tools like Terraform, Helm, Kustomize, and Ansible.
- Solid understanding of networking, security, and monitoring concepts in Kubernetes environments.
- Strong scripting and automation skills (Bash, Python, or similar).
- Excellent problem-solving and communication skills.
- Ability to work effectively in a fast-paced, collaborative environment.
Nice to Have:
- Certification in Kubernetes (CKA, CKAD) or related technologies.
- Experience with observability platforms like Prometheus, Grafana, Splunk, or OpenTelemetry.
- Knowledge of data processing frameworks like Apache Spark, Kafka, or Flink.
Location:
DBS Asia Hub
Job:
Analytics, Technology
Schedule:
Regular
Employee Status:
Full time
Similar roles
-
Site Reliability Engineer | Weekend Warrior
Jump Trading Amsterdam, North Holland, Netherlands · €150K–€175K/yr
-
[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Software Mind Krakow, Lesser Poland Voivodeship, Poland
-
Senior Site Reliability Engineer
Mozn Cairo, Cairo, Egypt
-
Manager- Site Reliability Engineering
Okta Bengaluru, Karnataka, India
-
Senior Site Reliability Engineer (12m FTC)
Mantel Sydney, New South Wales, Australia
-
Senior Site Reliability Engineer
2K Bangalore, Karnataka, India