Wayve

Senior Machine Learning Engineer, AI Performance

Wayve London, England, United Kingdom

Software Development · 1,001-5,000 employees

5 d ago
machine-learning Senior (5-10 yrs) Full-time United Kingdom
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the end-to-end delivery of model releases, from initial training and evaluation to final deployment readiness. You will also collaborate with cross-functional teams to optimize model performance and ensure systems meet strict runtime constraints.

What they look for

Machine Learning PyTorch Model Optimization Quantization Distillation TensorRT CUDA Qualcomm QNN Triton OpenCL Deep Learning Performance Engineering Latency Optimization Embedded Systems Analytical Skills Cross-functional Collaboration

Requirements

The role requires proven experience in improving production systems with tight constraints and strong hands-on experience training deep learning models in PyTorch. Proficiency with relevant toolchains like TensorRT or CUDA and the ability to reason across multiple levels of abstraction are essential.

Full description

The role

We’re looking for a Senior Machine Learning Engineer to join a high-ownership team responsible for delivering production-ready model releases as our OEM engagements and release cadence accelerate. This is an applied, delivery-focused MLE role—ideal for engineers who love shipping real systems and iterating quickly.

You’ll work on taking models from “works in training” to “meets product constraints,” partnering closely with teams downstream (e.g., inference/performance specialists) to ensure models are ready for deployment on-vehicle. As model capability grows, you’ll help keep the system within tight runtime constraints using a practical model optimisation techniques (e.g., quantisation, distillation, low-rank methods) where appropriate.

Key responsibilities

  • Own end-to-end delivery of model releases, from initial requirements through training, evaluation, iteration, and final readiness for deployment.
  • Train and iterate on PyTorch models with a strong experimental approach (hypothesis-driven iteration, ablations, clear evaluation criteria).
  • Debug and improve model performance using strong analytical skills—identifying regressions, root-causing issues, and proposing fixes.
  • Apply optimisation techniques (e.g., quantisation and distillation where beneficial), understanding trade-offs and when methods are appropriate.
  • Collaborate cross-functionally with adjacent ML and performance engineering teams to hand off models, define bottlenecks, and align on optimisation priorities.
  • Communicate clearly with stakeholders to align on delivery timelines, trade-offs, and readiness criteria.

About you

In order to set you up for success in this role at Wayve, we’re looking for the following skills and experience:

Essential

  • Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
  • Strong hands-on experience training and iterating on deep learning models in PyTorch (not just using high-level tooling).
  • Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly.
  • Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution.
  • Familiarity with model optimisation concepts such as quantisation and/or distillation (hands-on is a strong signal, but not a strict requirement if the fundamentals are solid).
  • Ability to reason across multiple levels of abstraction—from high-level model behaviour down to practical runtime/latency implications.
  • Strong engineering fundamentals and collaboration skills.

Desirable

  • Experience working on models that must meet tight latency / efficiency constraints (edge, embedded, real-time, or similarly constrained production settings).
  • Exposure to ML systems spanning training → evaluation → deployment handoff (even if you’re not writing kernels day-to-day).
  • Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.

#LI-HH1

Similar roles