About the role
You will design and build automated evaluation frameworks to measure the accuracy and reliability of AI-powered legal products. Additionally, you will collaborate with engineering teams to integrate these evaluations into development workflows and identify model failure modes.
What they look for
Requirements
The role requires a strong software engineering background with specific experience in AI/ML systems and evaluation methodologies. Proficiency in Python and the ability to translate complex legal workflows into technical evaluation criteria are essential.
Benefits
Full description
About Newcode.ai
Newcode.ai is a fast-growing legal tech and agentic AI company transforming how legal work is done. With teams across Norway, Sweden, Ireland, the US and growing we work at the intersection of law, technology, and intelligence. We move fast, think big, and take pride in doing things the right way.
Note: We believe in being transparent about what it's like to work at Newcode. As a fast-growing startup, we're building and evolving every day. That means not every process, playbook, or framework is already in place, and priorities can shift quickly.
The people who thrive here are comfortable with ambiguity, take ownership, and don't wait for perfect direction. They are resourceful, proactive, and able to "figure it out"—solving problems, creating structure where needed, and helping build the company as they go. If you are good with this then, great! Keep reading to learn more.
The Role
As a QA Engineer, you will own the systems, processes, and technical frameworks we use to evaluate the quality of Newcode's AI products.
You will work closely with Software Engineering and Legal Engineering to understand how our AI performs in real-world legal workflows, identify where it falls short, build better evaluation methodologies, and turn those insights into measurable improvements.
This is not traditional QA. You will be working at the intersection of AI evaluation, software engineering, legal workflows, and product quality.
Key Responsibilities:
- Design and build automated and human-in-the-loop evaluation frameworks for AI-powered products and workflows.
- Develop metrics and benchmarks to measure accuracy, reliability, consistency, and overall AI quality.
- Build evaluation datasets, test cases, and regression suites for complex legal use cases.
- Analyze model outputs to identify failure modes, patterns, and opportunities for improvement.
- Work closely with Legal Engineers to translate real-world legal workflows into rigorous evaluation criteria.
- Partner with Software Engineers to integrate evaluations directly into product and development workflows.
- Establish automated testing and monitoring to identify quality regressions before they reach customers.
- Evaluate changes to models, prompts, retrieval systems, agents, and workflows.
- Develop tooling that makes AI quality measurable and visible across the organization.
- Investigate unexpected model behavior and determine root causes.
- Help establish standards for what "good" looks like across different legal use cases.
- Strong software engineering or technical background, with experience working on AI/ML systems.
- Experience building testing, evaluation, or quality frameworks for AI-powered products.
- Strong Python and/or similar programming experience.
- Understanding of LLMs, generative AI, retrieval systems, agents, and/or AI evaluation methodologies.
- Strong analytical and problem-solving skills.
- Ability to work with ambiguous problems and turn them into measurable technical frameworks.
- Strong attention to detail and a high bar for quality.
- Ability to communicate technical findings clearly to both technical and non-technical stakeholders.
- Comfort working cross-functionally with Engineering, Product, and Legal Engineering.
Why You’ll Love Working with Us• Culture of trust and excellence: Be part of a team that values reliability, respect, and initiative, where everyone’s contribution matters and diversity drives creativity.
- Collaborative environment: Work with talented, driven peers on bleeding-edge legal engineering solutions with real-world impact.
- Growth and ownership: Shape your role as we grow. Take ownership and drive outcomes from the start of your journey.
- Inclusive culture of excellence and trust: We reward initiative, respect, integrity, and teamwork, and are committed to building a global, inclusive team. Your drive, curiosity, and ideas matter more than your background
Ready to shape what’s next? Apply today and join us in building a smarter, more efficient world powered by AI.
Similar roles
-
QA Compliance, Manager
Granules Pharmaceuticals Chantilly, Virginia, United States · $90K–$100K/yr
-
QA Engineer II (BENCH)
Arch Insurance Group Inc. Cebu, Isabela, Philippines
-
QA Engineer
Cresteo Provincia de Santiago, Santiago Metropolitan Region, Chile
-
QA Engineer
Fastbreak AI Charlotte, North Carolina, United States
-
Spécialiste en tests logiciel / Software Test Specialist - SC QA Domains
Genetec Montreal, Quebec, Canada
-
QA Production Holds Inspector
Steuben Foods Elma, New York, United States · $41K–$43K/yr