PwC

Data Scientist - AI Evaluation & Benchmarking Manager

PwC Birmingham, England, United Kingdom

Professional Services · 10,001+ employees

2 d ago
data-scientist Senior (5-10 yrs) Full-time Visa sponsorship United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and manage end-to-end AI benchmarking workflows and experimentation platforms to support evidence-based decision-making. Additionally, you will build scalable evaluation frameworks and pipelines while ensuring technical rigor and production-grade engineering.

What they look for

Data Science Python LLM experimentation Machine learning Cloud platforms CI/CD Containerization Experimental design Statistics Benchmarking Data engineering Asynchronous programming Technical leadership Model evaluation Infrastructure development Communication

Requirements

The role requires strong hands-on experience in data science and LLM experimentation using structured evaluation frameworks. Candidates must be proficient in Python and have experience deploying machine learning workloads to cloud platforms like Azure, AWS, or GCP.

Benefits

Private medical cover Virtual GP access Volunteering days Empowered flexibility

Full description

Line of Service

Internal Firm Services

Industry/Sector

Technology

Specialism

IFS - Information Technology (IT)

Management Level

Manager

Job Description & Summary

About the role:

You’ll join our AI Research team as a Data Scientist - AI Evaluation & Benchmarking Manager, helping to shape and evolve the model benchmarking and experimentation capabilities that underpin AI delivery across PwC and our clients. You’ll work within a highly collaborative applied research environment that values curiosity, technical rigour and practical problem solving. In this role, you'll develop frameworks and experimentation workflows used to evaluate emerging AI models, driving evidence-based decisions for client engagements. We’re looking for someone who enjoys technical ownership, thrives in fastmoving environments and is motivated by building scalable, secure AI research infrastructure.

What your days will look like:

You’ll play a key role in building and evolving our AI benchmarking and experimentation platforms, enabling robust and repeatable model evaluation. Your work will directly influence AI model selection and technical strategy across PwC projects and client engagements.

  • Design and run end to end benchmarking workflows, from understanding client use cases, designing and running benchmarking strategies, and generating business ready insights
  • Build scalable evaluation frameworks, metrics and pipelines, combining hands-on engineering with continuous review of academic literature to ensure our evaluation strategies reflect leading research and best practice.
  • Develop, maintain and improve experimentation infrastructure, ensuring robustness and production grade engineering.
  • Produce clear, client ready insights and support technical demos and deep dive sessions.

This role is for you if:

  • You have strong hands-on experience in Data Science concepts, or LLM experimentation using structured evaluation frameworks.
  • You are highly proficient in Python, including asynchronous programming, multithreading and writing maintainable code
  • You have experience deploying ML workloads to cloud platforms (Azure, AWS or GCP), with familiarity in CI/CD and containerisation (Docker/ Podman).
  • You have applied knowledge of statistics and experimental design and can translate findings into actionable recommendations. You are comfortable managing fastmoving workstreams and operating autonomously.
  • You demonstrate emerging leadership behaviours - taking initiative, influencing technical direction, communicating clearly and supporting the development of others. 
  • You are motivated by ownership and excited by the opportunity to shape AI research platforms that directly impact client engagements.

What you’ll receive from us:

No matter where you may be in your career or personal life, our benefits are designed to add value and support, recognising and rewarding you fairly for your contributions.

We offer a range of benefits including empowered flexibility and a working week split between office, home and client site; private medical cover and 24/7 access to a qualified virtual GP; six volunteering days a year and much more.

 

Education (if blank, degree and/or field of study not specified)

Degrees/Field of Study required:

Degrees/Field of Study preferred:

Certifications (if blank, certifications not specified)

Required Skills

Optional Skills

Accepting Feedback, Accepting Feedback, Active Listening, Analytical Thinking, Artificial Intelligence, Big Data, C++ Programming Language, Coaching and Feedback, Communication, Complex Data Analysis, Creativity, Data-Driven Decision Making (DIDM), Data Engineering, Data Lake, Data Mining, Data Modeling, Data Pipeline, Data Quality, Data Science, Data Science Algorithms, Data Science Troubleshooting, Data Science Workflows, Deep Learning, Embracing Change, Emotional Regulation {+ 22 more}

Desired Languages (If blank, desired languages not specified)

Travel Requirements

Up to 20%

Available for Work Visa Sponsorship?

Yes

Government Clearance Required?

No

Job Posting End Date

August 21, 2026

Similar roles