Babbel

Senior Machine Learning Engineer (all genders) - Babbel Labs

Babbel Berlin, Germany

E-Learning Providers · 501-1,000 employees

20 h ago
Remote machine-learning Senior (5-10 yrs) Full-time Germany
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will own and improve the learner-personalisation engine by taking designs from specification to production. This involves rigorous benchmarking, monitoring system performance, and conducting online experiments to validate model improvements.

What they look for

Machine Learning Python TypeScript Probabilistic Modeling Bayesian Inference AWS Terraform CI/CD Recommendation Systems Ranking Systems Scoring Systems A/B Testing Data Analysis Model Evaluation Graph ML

Requirements

The role requires strong hands-on experience in shipping machine learning models to production and proficiency in Python or TypeScript. Candidates should have a solid understanding of probabilistic modeling, system evaluation, and modern cloud infrastructure tools.

Full description

Senior Machine Learning Engineer (all genders) - Babbel Labs

About Babbel Labs

Babbel Labs is building the future of language learning. We are an AI-first, independent company within the Babbel group, based in Berlin. Our teams bring AI, research and product together to ship experiences that set a new standard for how people learn.

The role

Babbel's learner-personalisation engine tracks what a learner has and hasn't mastered, and decides what they should practice next. We are actively pushing personalisation further; moving beyond describing a user and prescribing targeted practice, to adapting the learning journey ahead of them.

This is a hands-on senior individual-contributor role on that team. You will own real subsystems end to end and make decisions about how to improve the personalisation engine. That includes everything from research and benchmarking through production, monitoring and incidents that follow six months later.

How you'll make an impact

  • Work directly with the Principal Scientist to take designs from spec into production, then keep improving what you’ve built on your own judgement rather than waiting to be told what’s next.
  • Help shape new features from the beginning - not just implementation of a spec handed to you.
  • Take real ownership of core personalisation and mastery-tracking subsystems, operate independently, and make decisions confidently and competently.
  • Design the evaluation that tells you whether a change is real: a rigorous offline benchmark against a real baseline, and the online experiment that confirms or kills it.
  • Run what you build. Instrument it, notice when it's silently wrong rather than only when it errors, and fix it before it becomes an incident.
  • Deliver with coding agents as a matter of course, and verify what they produce before you rely on it.

Your skills and qualifications

  • Strong, hands-on ML engineering that has shipped real models to production — recommendation, ranking, scoring, or trust-and-safety systems under real user load are the closest match. Research or competition experience is a plus.
  • Experience with probabilistic modeling, latent-variable modeling and Bayesian inference, or the equivalent rigor from an adjacent domain.
  • ML system evaluation: monitoring metrics you define, debugging output that doesn’t look right, and rolling out a change to a live scoring or ranking system without breaking it.
  • Rigorous experimentation practice: benchmarking against a real baseline, running or reading A/B tests correctly, and the judgement to know when an offline improvement won't survive contact with production.
  • TypeScript/Python as your primary languages, with enough command of our surrounding stack (AWS, Terraform, CI/CD) to ship and own your own service's delivery. This is not an infrastructure role, so depth there is not the bar.
  • Coding agents are part of your daily workflow, and you check their output before you rely on it. You are neither dismissive of them nor careless with them.

Nice to have

  • Experience with psychometric models, such as Item Response Theory.
  • Graph ML experience — embeddings, graph neural networks, or relational modeling — at real scale.
  • Public technical work: open-source contributions, writing, or competitive ML.

Ways of working

This is a hands-on individual-contributor role on a small team that moves fast, agent-first, but with a disciplined delivery approach. It is remote and Berlin-friendly. The interview process is deliberately short: an initial screen, followed by one structured working session with the hiring manager.

Diversity at Babbel

As part of our ongoing journey towards building a diverse, equitable, and inclusive company, we welcome everyone to apply, especially individuals who are underrepresented in tech. We are a learning company, inside and out, and we encourage you to apply even if you do not fit all the technical requirements — all candidates are assessed based on skills, qualifications, and our business needs. Please state your pronouns in your application, and let us know if you'd like to be addressed by a name other than the one appearing on your official documents. If you have a disability or special need, feel welcome to inform us so we can provide proper assistance in the application process.

Similar roles