Support Engineer – AI Server Infrastructure
systemsGo Tokyo, Tokyo, Japan
IT Services and IT Consulting · 51-200 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
The Support Engineer will perform preventative and corrective maintenance on AI server systems and associated infrastructure across Japan. They are responsible for diagnosing hardware failures, managing system deployments, and providing professional technical support to enterprise customers.
What they look for
Requirements
Candidates must have at least 3 years of experience in enterprise server infrastructure and hands-on hardware troubleshooting. Proficiency in Linux, basic networking, and business-level Japanese language skills are required.
Full description
Location: Tokyo, Japan
About the Opportunity
Our client is a global technology company at the forefront of next-generation AI computing and high-performance infrastructure solutions. The organization develops cutting-edge hardware and software technologies that power advanced artificial intelligence, machine learning, and large-scale computing environments.
With a strong engineering-driven culture and an expanding international presence, the company is committed to delivering scalable, reliable, and high-performance infrastructure solutions for enterprise and research customers worldwide.
Position Summary
We are seeking a skilled and customer-focused Support Engineer – AI Server Infrastructure to support the deployment, maintenance, and operation of AI server systems across Japan.
The successful candidate will provide both on-site and remote technical support for high-performance servers, storage platforms, networking equipment, and related infrastructure. This role requires strong troubleshooting skills, hands-on hardware experience, and the ability to collaborate effectively with customers and global engineering teams.
Key Responsibilities
AI Infrastructure Support
- Perform preventative and corrective maintenance on AI server systems and associated infrastructure.
- Diagnose and resolve hardware failures involving servers, accelerators, memory, storage, power supplies, and PCIe devices.
- Conduct component replacements and on-site repair activities.
- Support server deployments, installations, relocations, and decommissioning projects.
- Execute firmware, BIOS, driver, and software upgrades.
Monitoring & Incident Response
- Monitor system availability and performance using remote management and monitoring tools.
- Investigate and analyze system alerts, logs, and hardware events.
- Perform first-level troubleshooting and root cause identification.
- Create incident reports and maintain accurate service documentation.
- Escalate complex technical issues to engineering and global support teams as required.
Customer Engagement
- Deliver professional technical support at customer offices and data center facilities.
- Communicate technical findings, risks, and resolutions effectively to customers and internal stakeholders.
- Ensure support activities are completed in accordance with service-level agreements (SLAs).
- Act as a trusted technical advisor during maintenance windows and critical outage situations.
Asset & Inventory Management
- Track and manage spare parts inventory.
- Coordinate replacement part shipments and logistics activities.
- Maintain maintenance reports, service records, and asset tracking documentation.
Required Qualifications
- Minimum 3 years of experience supporting enterprise server infrastructure.
- Hands-on experience troubleshooting x86 servers and enterprise hardware platforms.
- Strong understanding of hardware diagnostics and failure isolation.
- Working knowledge of Linux operating systems (Ubuntu, RHEL, CentOS, etc.).
- Basic networking knowledge, including:
- TCP/IP
- DHCP
- VLANs
- Layer 2 / Layer 3 networking
- IPMI
- Experience providing customer-facing technical support and on-site services.
- Familiarity with hardware diagnostic tools such as:
- ipmitool
- smartctl
- nvidia-smi (or equivalent monitoring tools)
- Experience preparing technical and incident documentation.
- Ability to read and understand English technical documentation.
- Business-level Japanese language skills.
- Conversational English communication ability.
Preferred Qualifications
- Experience supporting GPU servers, AI infrastructure, or high-performance computing (HPC) environments.
- Hands-on experience with enterprise AI server platforms and large-scale compute clusters.
- Knowledge of:
- InfiniBand
- NVLink
- PCIe architectures
- High-speed Ethernet networking
- Experience working within data center environments.
- Understanding of AI, machine learning, or HPC workloads.
- Experience with Linux shell scripting and automation.
- Valid Japanese driver's license.
- Experience supporting multinational enterprise customers.
Language
- Japanese: Business
- English: Business level
Similar roles
-
Consultant Systems Support Engineer
Thoughtworks Bucharest, Romania
-
Lead Application Support Engineer (6AM to 3PM Shift)
DTCC Makati, Metro Manila, Philippines
-
IT Support Engineer
AEGEAN Athens, Attica, Greece
-
Voice Services Support Engineer (Cisco CUCM / OpenScape Alarm Response)
Spektrum Mons, Wallonia, Belgium
-
IT Support Engineer (3 month FTC)
Oxford Quantum Circuits Reading, England, United Kingdom
-
Technical Support Engineer
Telana Preston, England, United Kingdom