Machine Learning Inference Manager
Stats Perform · Prague, Prague, Czechia
Spectator Sports · 1,001-5,000 employees
About the role
The Machine Learning Inference Manager will own the strategy, architecture, and operation of the model serving platform across on-premise and cloud environments. This role involves leading a team of engineers to ensure reliable, low-latency, and cost-efficient production inference for various machine learning workloads.
What they look for
Requirements
Candidates must have a strong software engineering background with production experience in deploying and operating ML models at scale. Proficiency in GPU-based inference, containerization, orchestration, and cloud infrastructure is essential for this leadership position.
Benefits
Full description
Overview
Stats Perform is the market leader in sports tech. We provide the most trusted sports data to some of the world's biggest organizations, across sports, media, and broadcasting.
Through the latest AI technologies and machine learning, we combine decades' worth of data with the latest in-game happenings. We then offer coaches, teams, professional bodies, and media channels around the world, access to the very best data, content, and insights. In turn, improving how sports fans interact with their favourite sports teams and competitions.
How do we add value?
- Media outlets add a little magic to their coverage with our stats and graphics packages.
- Sportsbooks can offer better predictions and more accurate odds.
- The world's top coaches are known to use our data to make critical team decisions.
- Sports commentators can engage with fans on a deeper level, using our stories and insights.
Anywhere you find sport, Stats Perform is there. However, data and tech are only half of the package. We need great people to fuel the engine.
We succeeded thanks to a team of amazing people. They spend their days collecting, analyzing, and interpreting data from a wide range of live sporting events. If you combine this real-time data with our 40-year-old archives, elite journalists, camera operators, copywriters, the latest in AI wizardry, and a host of 'behind the scenes' support staff, you've got all the ingredients to make it a magical experience!
Responsibilities:
We are looking for a Machine Learning Inference Manager to own the strategy, architecture, and day-to-day operation of our model serving platform. This is the person accountable for how models move from a trained artefact into reliable, low-latency, cost-efficient production inference- across real-time streaming, near-real-time, and batch workloads, on both on-premise GPU infrastructure and cloud.
The role blends hands-on technical leadership with line management. You will set the technical direction for inference serving, define and defend performance and reliability SLOs, optimise GPU utilisation and unit economics, and grow a small team of ML platform/ inference engineers. You will act as the bridge between the research and applied-ML teams who produce models and the product and operations teams who depend on them being fast, available, and affordable.
Key Responsibilities
Inference Platform Ownership
- Own the end-to-end inference serving stack: model packaging, deployment, versioning, routing, autoscaling, and decommissioning.
- Define and maintain the reference architecture for serving models across on-premise GPU clusters and cloud, including the hybrid boundary between the two.
- Establish standards and paved-road tooling so that model teams can deploy safely and repeatably without deep infrastructure knowledge.
Performance, Latency & Optimisation
- Set, measure, and enforce latency, throughput, and availability SLOs for each class of inference workload (real-time, streaming, batch).
- Drive inference optimisation: batching strategies, quantisation, distillation-aware serving, kernel/runtime selection, and hardware-aware tuning.
- Lead GPU efficiency initiatives - utilisation, memory management, multi-model serving, and right-sizing - to maximise throughput per unit of hardware spend.
Reliability & Observability
- Own the reliability posture of the serving platform: monitoring, alerting, capacity planning, failover, and incident response for inference services.
- Implement observability for model latency, throughput, error rates, drift signals, and cost per inference, with clear dashboards for both engineering and leadership.
- Establish rollout and rollback practices (canary, shadow, blue/green) so model updates ship safely.
Cost & Capacity Management
- Own the inference cost model and forecast; report unit economics (e.g. cost per lk inferences) and drive them down over time.
- Plan GPU capacity across on-prem and cloud, balancing performance, resilience, and spend.
Team Leadership
- Line-manage and mentor a team of inference / ML platform engineers, setting objectives, running the delivery cadence, and supporting career growth.
- Partner with research, applied ML, data engineering, product, and SRE to align the serving roadmap with business priorities.
- Contribute to hiring, onboarding, and building an engineering culture grounded in reliability, security, and pragmatism.
Governance & Security
- Ensure the inference platform meets security, data-protection, and compliance requirements, including appropriate handling of any personal or sensitive data flowing through inference.
- Maintain clear model lineage, access controls, and audit trails for deployed models.
Required Qualifications:
- Strong software engineering background with production experience deploying and operating ML models at scale.
- Deep, hands-on knowledge of at least one model-serving framework (e.g. NVIDIA Triton Inference Server, TorchServe, vLLM, TensorRT-LLM, Ray Serve, or KServe).
- Demonstrable expertise in GPU-based inference: performance profiling, batching, quantisation, and runtime optimisation (e.g. TensorRT, ONNX Runtime).
- Solid experience with containerisation and orchestration (Docker/Podman, Kubernetes) in production.
- Cloud experience (AWS preferred), including GPU compute, autoscaling, and cost management.
- Proficiency in Python, with the ability to read and reason about model code and serving internals.
- Proven experience defining and operating against SLOs, with strong instincts for observability and incident response.
- Experience leading or mentoring engineers, whether as a line manager or a senior technical lead ready to step up.
Desired Qualifications:
- Experience serving models under strict real-time or streaming latency constraints (e.g. video, audio, live analytics).
- Familiarity with hybrid on-premise+ cloud inference topologies and the trade-offs involved.
- Experience with LLM serving and optimisation (KV-cache management, continuous batching, speculative decoding, quantised inference).
- Exposure to computer-vision or sequence-model inference pipelines in production.
- Background in cost/performance optimisation oflarge GPU fleets.
- Experience with CI/CD for ML (model registries, automated deployment pipelines) and infrastructure-as-code.
Why work at Stats Perform?
We love sports, but we love diverse thinking more!
We know that diversity brings creativity, so we invite people from all backgrounds to join us. At Stats Perform you can make a difference, by using your skills and experience every day, you'll feel valued and respected for your contribution.
We take care of our colleagues
We like happy and healthy colleagues. You will benefit from things like Mental Health Days Off, ‘No Meeting Fridays,’ and flexible working schedules.
We pull together to build a better workplace and world for all.
We encourage employees to take part in charitable activities, utilize their 2 days of Volunteering Time Off, support our environmental efforts, and be actively involved in Employee Resource Groups.
Diversity, Equity, and Inclusion at Stats Perform
By joining Stats Perform, you'll be part of a team that celebrates diversity. A team that is dedicated to creating an inclusive atmosphere where everyone feels valued and welcome. All employees are collectively responsible for developing and maintaining an inclusive environment. That is why our Diversity, Equity, and Inclusion goals underpin our core values.
With increased diversity comes increased innovation and creativity. Ensuring we're best placed to serve our clients and communities. Stats Perform is committed to seeking diversity, equity, and inclusion in all we do.