P

AI - QA Engineer

Publicis Groupe Holdings B.V Mumbai City District, Maharashtra, India

Information Services · 11-50 employees

3 h ago
qa Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Define and implement quality standards and release gates for AI agents to ensure accuracy, security, and performance. Build golden datasets and conduct UAT cycles to validate agent behavior, citation grounding, and adversarial robustness.

What they look for

QA Engineering LLM Testing RAG Agent Systems Automated Testing CI/CD Prompt Injection Testing Data Leakage Prevention Citation Accuracy Hallucination Detection UAT Golden Datasets Security Trimming Tool Selection Regression Testing

Requirements

Requires 5-8 years of experience in QA engineering with at least 2 years specifically focused on LLM, RAG, or agent systems. Candidates must demonstrate the ability to build automated evaluation suites and translate business feedback into repeatable test cases.

Full description

Overview

Mission

Own release readiness for agents whose outputs go into client-facing RFPs, data grids and pitch decks. Test that answers are grounded and correctly cited, that users never see content they are not entitled to, that multi-agent handoffs behave, and that generated files are correct. Then turn this into automated gates that every agent must pass.

Responsibilities

Key Responsibilities

  • Define quality standards and release gates per agent (groundedness, citation accuracy, hallucination, task completion, latency, cost) aligned to the PRD's decision-grade vs good-enough split.
  • Build golden datasets with Global BD SMEs from real RFPs, grids and prior submissions; run UAT cycles with named business testers.
  • Test citations by checking every claim links to a real, permission-appropriate source, and that gaps and contradictions are flagged rather than invented.
  • Test permissions using personas with different access to prove security trimming holds across retrieval and live tool calls.
  • Test agent behaviour including tool selection, execution correctness, failure handling and multi-agent handoffs / downstream triggers.
  • Test document outputs for grid cell accuracy, template fidelity, and tracked-change style proofing suggestions rather than silent edits.
  • Run adversarial tests for prompt injection (including via uploaded RFP files), data leakage and safety boundaries.
  • Automate regression for model, prompt, embedding and retrieval changes inside the CI/CD pipeline.

Qualifications

Experience & Working Style

  • 5–8 years in QA / test engineering, with 2+ years testing LLM, RAG or agent systems.
  • Has built an automated LLM evaluation suite that gated real releases.
  • Can sit with business reviewers, capture why an output is wrong and turn it into a repeatable test.

Not required (don't screen out for these): Traditional manual-only QA backgrounds are not sufficient. Deep ML modelling knowledge is not required.

Similar roles