NVIDIA

Senior Data Platform Software Engineer

NVIDIA · Tel Aviv, Tel-Aviv District, Israel

Computer Hardware Manufacturing · 10,001+ employees

19 h ago
Senior (5-10 yrs) Full-time Israel
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Lead the development of advanced metric and measurement tools for real-time data collection and processing across NVIDIA's AI hardware components. Collaborate with applied researchers, architects, and data engineers to iterate on ML diagnostic toolkits and improve hardware visibility.

What they look for

Data acquisition Predictive maintenance Root-cause analysis AIOPS Metric extraction Telemetry Hardware diagnostics Machine learning Statistics Networking System architecture Performance behavior Failure analysis Data engineering AI hardware

Requirements

Requires a Bachelor's degree in Computer Science, Electrical Engineering, or a related field with at least 5 years of hands-on experience. Candidates must possess deep system knowledge, strong networking and hardware/software understanding, and proficiency in AI/ML modeling and statistics.

Full description

We are seeking a highly skilled Data Platform SW Engineer to join the Applied Networking AI group. In this role you will help develop advanced data acquisition solutions for the fields of predictive-maintenance, root-cause analysis and AIOPS. You'll collaborate closely with subject-matter-experts (SMEs), applied-researchers, architects, data-engineers and other stakeholders to push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.

As a key contributor you will develop and own metric-extraction, measurement and telemetry tools that enable a high resolution viewpoint into the hardware. You will experiment and iterate fast and in collaboration with applied-researchers to improve our ML diagnostic and prediction toolkit.

What you'll be doing:

  • Lead the development of advanced metric and measurement tools for real-time data collection and processing, to enable a high resolution viewpoint into the full set of HW components that compose NVIDIA's AI factory solutions (GPUs, networking interfaces, etc).
  • Work alongside applied-researchers to experiment and iterate on the bridge between metrics and ML.
  • Partner with architects and product managers to gain a deep understanding of NVIDIA’s hardware and roadmap.
  • Collaborate with data-engineers to enable high resolution tools at scale.

What we need to see:

  • BSc in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
  • 5+ years of hands-on experience demonstrating deep system knowledge and metric extraction development.
  • Strong understanding of networking, hardware/software systems, performance behavior, or failure analysis.
  • Solid understanding of AI/ML modeling and statistics.
  • Excellent ability to convey and communicate data-based insights to stakeholders and management.

Ways to stand out from the crowd:

  • Experience diagnosing or debugging modern AI hardware.
  • High energy and a positive, proactive and curious approach.

We are an equal opportunity employer and value diversity at our company. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.