Senior Data Scientist, AI Scoring & Evaluation
Workera AI East Flanders, Flanders, Belgium
Technology, Information and Internet · 51-200 employees
About the role
You will own the end-to-end scoring quality of the assessment platform, including evaluator design, calibration, and accuracy against expert benchmarks. You will also lead analytics investigations and serve as the liaison between AI governance and InfoSec to ensure defensible and fair measurement.
What they look for
Requirements
The role requires 3-4+ years of production data science experience, specifically with Python and SQL, and a strong background in statistical analysis. Experience with LLM-based systems and a desire to learn psychometrics or educational measurement are highly valued.
Full description
Senior Data Scientist - Assessment Scoring & Evaluation
Everyone's racing to build AI. Workera exists for the 8 billion people who work alongside it.
While the world's attention is on creating new tools, someone has to solve the other side of the equation: the humans. The workforce is going through the biggest transformation in a generation and most organizations are navigating it blind, without the data to understand what their people can actually do, where the gaps are, or how to close them fast enough.
As Workera moves into higher-stakes decisions, you're the person who can prove a measurement is accurate, fair, and defensible. You'll own the measurement system that turns signals into the skill data our customers act on: the evaluation harnesses, the calibration methods, the quality monitoring behind every score. You'll also shape how our AI software itself is designed, from how components are orchestrated and prompted to the evaluation frameworks that decide whether they're good enough to ship.
You're not just running analyses, you're building the trust layer underneath a product that increasingly makes high-stakes calls about people's careers. If you want a role where your judgment becomes the thing customers, auditors, and your own engineering team rely on, and where the measurement problems are still being invented rather than optimized, this is it.
Workera's skills intelligence platform is critical infrastructure for the AI era: the layer that lets organizations understand, mobilize, manage, and develop their talent with precision. We're trusted by the Fortune 500, powered by proprietary AI agents, and built by a small, senior team, which means what you ship here has outsized reach.
WHY THIS ROLE EXISTS
Workera is scaling fast: more customers, more use cases, more fields evaluated, higher stakes on every measurement. That scale changes what quality means. What we once crafted and inspected by hand now needs monitoring that catches issues before customers do, and remediation that resolves them fast.
This role sits at the intersection of Engineering, Assessment Science, and Design, partnering closely with the product engineers who build our scoring pipeline. You're the data layer that pieces those disciplines together: the foundation the rest of our scoring is built on, and the reason a construct definition from Assessment Science actually turns into a number a customer can trust.
YOUR TEAM
You report to Dr. Taylor Sullivan, Workera’s VP of Product and Assessments, and work day to day with two groups: The assessment science team decides what we're measuring and how a skill turns into test content, they own construct definition (what "good at X" actually means), blueprinting (how an assessment is structured), and content authoring (writing the actual questions). You partner with them on how scoring methods serve that intent. The assessment tech team builds and maintains the scoring pipeline; you work embedded with them: writing specs, implementing and coordinating improvements, and validating that what ships meets the quality bar. You also work with data engineering on data availability. It's a small, cross-functional team, and you sit inside product decisions rather than in a separate research function.
WHAT YOU'LL OWN
Build trust by owning our scoring mechanism. These are the outcomes you're accountable for:
- Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability for accuracy against expert benchmarks
- Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and the pre-release gate that every scoring change must clear
- Define and publish assessment quality KPIs (human to AI agreement, reliability, classification accuracy, bias indicators, latency, cost per assessment) on dashboards anyone in the company can reference
- Lead the analytics investigations behind assessment decisions: performance studies, impact simulations, root cause analysis when scores look wrong
- Ship measurement improvements end to end: implement and coordinate with tech, own the spec, the validation, and the quality bar in both cases
- Be Workera's liaison between AI governance and InfoSec: stay current on frameworks like GDPR and the EU AI Act, and turn that into the evidence we show in audits, security reviews, and enterprise diligence when customers ask how we use AI in our assessments
- Empower partners across Product, Engineering, and GTM to get the data and analysis they need autonomously, by building AI tooling and documentation rather than answering each request yourself
HOW YOU'LL RAMP
We don't expect you to figure it out alone. Here's what great looks like at each stage:
First 30 Days: Learn the Machine
- Immerse yourself in Workera's platform, customers, and the problems we're solving. You'll shadow key workflows and understand how AI is embedded in day-to-day operations across teams
- Deliver at least one analytics investigation that changes a product decision
- Sit in on an enterprise or auditor conversation about score quality to see what evidence customers actually ask for
By 60 Days: Ship Something Real
- Own your first meaningful deliverable and demonstrate end-to-end execution
- Troubleshoot scoring end to end independently
- Ship one measurable scoring improvement, validated against expert labelled data, through the pre-release gate
By 90 Days: Multiply Your Impact
- Operate with full autonomy in your domain; your team relies on your judgment
- Have built or deployed at least one AI-assisted workflow that the team adopts
- Set the scoring roadmap: identify the gaps, prioritise what gets built, and drive those initiatives to completion with Engineering
- Build systems that improve themselves: evaluation loops that flag their own drift, calibration that updates on new expert labels, tooling that gets better as it is used
We're a fast-moving company -- the scope and shape of this role will evolve as we do.
WHAT YOU BRING
We're looking for signal, not checkboxes. In rough order of what matters most:
- You've driven production data science work independently.
- 4+ years (or 3+ with demonstrated end-to-end ownership) of a production data project ideally in an AI startup or fast-moving product environment where you scoped investigations yourself rather than picking up defined tickets.
- Python and SQL are assumed; you shouldn't need an engineer to run an analysis.
- Real statistical and analytical depth. You can turn, "is our scoring reliable?" into a defined study with a defensible method and a clear answer fast, and without being handed the design.
- You can read this kind of data. Quantitative work with educational, learning, or assessment data or a clear appetite to go deep on it fast. We're looking for someone who can interpret our assessment data fluently and wants to become expert in how it's built.
- You've worked with LLM-based systems in production. Prompt design, evaluating model output against human judgment, and a real sense of where these systems fail.
- You explain quantitative results to non-technical audiences, including executives and customers, without losing the nuance.
- A plus, not a requirement: background in psychometrics or educational measurement. Our assessment scientists will coach you — but you should want to learn it.
HOW WE WORK: AI IS THE DEFAULT
At Workera, AI isn't a feature we sell -- it's how we operate. Every team member is expected to:
- Use AI daily. AI assistants, copilots, and automation tools are part of your stack -- not optional extras. We expect you to actively experiment with new tools and push the boundary of what's possible in your function.
- Build your own leverage. Our marketers write code. Our PMs build automations. Our ops team deploys agents. If a workflow can be automated, you're expected to automate it.
- Think in systems, not tasks. We value people who build repeatable, scalable solutions over people who grind through one-off work. Your goal is to make your function run smarter, not just harder.
AI fluency is a cultural expectation, not a line item on a job description.
ABOUT WORKERA
We're a Silicon Valley company backed by NEA, Jump Capital, and Owl Ventures. Our founder is Kian Katanforoosh, an award-winning Stanford Computer Science Lecturer who has taught AI to over 1 million people. Our Chairman is Dr. Andrew Ng, co-founder of Coursera, CEO of DeepLearning.AI, and founding lead of the Google Brain project.
Our clients include Accenture, Siemens Energy, Samsung, and the United States Air Force.
Named to Fast Company's Most Innovative Companies list alongside Microsoft and Canva. Recognized by the World Economic Forum's Tech Pioneers, Inc 5000, and Josh Bersin's HR Tech AI Trailblazers. In a world where every company claims to 'do AI', at Workera, it's actually in our DNA.
We're learners, builders, and dreamers. Join us.
Workera is committed to providing an inclusive and respectful environment where equal employment opportunities are available to all applicants and employees. We do not discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law. Hiring decisions are based on qualifications, merit, mindset, and business need.
About Workera We're a Silicon Valley company backed by NEA, Jump Capital, and Owl Ventures. Our founder is Kian Katanforoosh, an award-winning Stanford Computer Science Lecturer who has taught AI to over 1 million people. Our Chairman is Dr. Andrew Ng, co-founder of Coursera, CEO of DeepLearning.AI, and founding lead of the Google Brain project.
Our clients include Accenture, Siemens Energy, Samsung, and the United States Air Force.
Named to Fast Company's Most Innovative Companies list alongside Microsoft and Canva. Recognized by the World Economic Forum's Tech Pioneers, Inc 5000, and Josh Bersin's HR Tech AI Trailblazers. In a world where every company claims to 'do AI', at Workera, it's actually in our DNA.
We're learners, builders, and dreamers. Join us. Workera is committed to providing an inclusive and respectful environment where equal employment opportunities are available to all applicants and employees. We do not discriminate on the basis of race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, veteran status, or any other characteristic protected by applicable law. Hiring decisions are based on qualifications, merit, mindset, and business need.
Similar roles
-
Data Scientist
EWC Corporate LLC Plano, Texas, United States
-
Dairy Data Scientist
Alta Genetics United States
-
Senior Data Scientist
Midcontinent Independent System Operator (MISO) Carmel, Indiana, United States · $119K–$144K/yr
-
Data Scientist
Signature Aviation Orlando, Florida, United States
-
Data Scientist Principal
ADT Irving, Texas, United States
-
Data Scientist
ADT Irving, Texas, United States