Oxydata Software

Senior Data Centre Operations Engineer

Oxydata Software · Senai, Johor, Malaysia

IT Services and IT Consulting · 51-200 employees

4 h ago
Senior (5-10 yrs) Full-time Malaysia
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The engineer will manage and support server and GPU infrastructure, including hardware configuration, troubleshooting, and performance monitoring. They are also responsible for automating routine tasks, maintaining technical documentation, and coordinating with vendors for hardware maintenance.

What they look for

Server hardware troubleshooting GPU infrastructure management Linux administration NVIDIA-SMI DCGM BIOS management BMC management RAID configuration Firmware upgrades Hardware alert management Technical documentation Vendor coordination Power and cooling monitoring Shell scripting Ansible Network topology

Requirements

Candidates must have a bachelor's degree in a relevant field and 5-7 years of experience in server operations. Strong Linux administration skills, hardware troubleshooting expertise, and proficiency with NVIDIA diagnostic tools are required.

Full description

Senior Data Centre Operations Engineer

Location: Senai, Johor, Malaysia

Work Mode: Onsite

Employment type: Permanent

Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.

We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.

Responsibilities

  • Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
  • Monitor server health and review IPMI, BMC, and operating system logs.
  • Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
  • Manage RAID configurations and monitor SSD and NVMe health.
  • Troubleshoot GPU servers and replace faulty hardware components.
  • Support GPU cluster performance, network topology, and system stability.
  • Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
  • Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
  • Prepare technical guides, troubleshooting documentation, and standard operating procedures.
  • Coordinate with hardware vendors and manage RMAs and spare parts.
  • Monitor rack power, temperature, and air-cooling conditions.

Requirements

Must-have:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
  • Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
  • At least 5-7 years of experience in server operations.
  • Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
  • Strong Linux administration and troubleshooting knowledge.
  • Good hardware troubleshooting and problem-solving skills.
  • Ability to work effectively with internal teams and external vendors.
  • Good technical communication skills in English.

Nice-to-have:

  • Experience operating large-scale GPU clusters.
  • Experience with Shell or Ansible automation.
  • Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
  • Exposure to air-cooled server environments.
  • Experience with liquid-cooled servers.
  • Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
  • Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.

Education:

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.

Why Join Us

  • Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
  • Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.

Apply Now: https://www.careers-page.com/oxy/job/5W8V6Y35