DBS Bank

Senior Associate, Observability Platform SRE Engineer, SRE & Governance, Group Technology

DBS Bank · Singapore, Singapore, Singapore

Banking · 10,001+ employees

13 h ago
Mid (2-5 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The SRE engineer will ensure high availability and resiliency of platform services while automating routine tasks to improve team productivity. They are responsible for implementing observability standards, managing monitoring solutions, and resolving performance issues through detailed system analysis.

What they look for

Site Reliability Engineering Observability ELK Stack Grafana Open Telemetry Automation Scripting Performance Management Capacity Planning CI/CD DevOps Kafka Prometheus Network Routing Load Balancing Troubleshooting

Requirements

Candidates must have at least 3 years of IT experience and a university degree in computer science or a related field. Strong knowledge of SRE practices, observability tools, and scripting languages is required for this role.

Full description

Job Objective

DBS Bank is looking for a Platform SRE Engineer with experience working on enterprise level data engineering, analytics, and observability applications. The SRE engineer would be responsible for ensuring high availability of the platform services and perform continuous improvements to increase the platform’s efficiency and resiliency. The SRE engineer will also perform automation development tasks to remove toil and increase the team’s productivity.

Roles and Responsibilities

Develop monitoring and onboarding guidelines for various applications using observability platform stack, ensuring accurate monitoring and data collection.

Implement Observability standards, best practices, operations and processes for the Enterprise in Observability tools.

Automate routine tasks and reporting processes using APIs and scripting, reducing manual effort and improving efficiency in AppDynamics & other observability tools

Identify and resolve performance issues through detailed analysis of transaction traces, application logs, and system metrics.

Contribute to internal knowledge bases, create documentation, and share insights with the team to promote a culture of learning and collaboration.

Design and implement monitoring solutions to track application performance, identifying bottlenecks, capacity planning and optimising system efficiency.

Develop custom dashboards and reports to provide actionable insights

Collaborate with development and operations teams to integrate Observability platform stack with CI/CD pipelines and other DevOps tools.

Configure and fine-tune alerts to proactively detect and address performance issues before they impact end-users.

Create data retention polices and access controls (RBAC) to manage user permissions.

Perform application maintenance, patching, upgrading controller versions, agents etc and ensure EOS/EOL is maintained.

Deliverables

Ensure on-time delivery of tasks and projects.

Ensure continuous uptime of applications and services.

Ensure no security or audit issues

Job Dimensions

Comply to bank standards to track and follow up on the assigned projects.

Cover all areas in application and infrastructure operations of the platform.

Education and Relevant Experience

You should be a university graduate (computer science or related field) with good experience working with contemporary technologies and scripting languages.

Strong communication skills and ability to explain protocol and processes with team and management

A passion for learning and using new technologies in the open-source communities.

Functional / Technical Competencies

Min 3 years of IT work experience.

Working knowledge in ELK Stack, Grafana, Open Telemetry (OTEL)

Experience in triaging and troubleshooting application problems quickly in monitoring tools by using various techniques

Knowledgeable and experienced in SRE (Site Reliability Engineering) practices covering monitoring, observability, performance management, automation, and resiliency.

Knowledge in Confluent Kafka, Prometheus & other APM tools (Dynatrace, Datadog, New Relic, Splunk) is a plus.

Knowledge in AI/ML capabilities to automate RCA’s and shorter MTTR when issues arise.

Good understanding of Network routing, Load balancing and Networking protocols; a base knowledge of TCP/IP, with an understanding of HTTP and DNS

Good problem diagnosis and creative problem-solving skills

Location:

DBS Asia Hub

Job:

Technology

Schedule:

Regular

Employee Status:

Full time