Programme at a glance

A large language model builder needed structured evaluation and A and B testing across several models at once, covering both general use and domains where a wrong answer carries real consequences. The requirement was continuous rather than one off. Every sprint fed the next training cycle, so the evaluation had to arrive at the pace the model was changing.

Evaluations delivered to date420,000+
Sprint cadence40,000+ evaluations per 6 day sprint
Models compared per round4 to 6, side by side
Domains covered8, including clinical, legal, financial services, software engineering, mathematics and automotive
Evaluator pool180 subject matter experts, credentialed in their field
New domain stand up9 days from brief to first scored batch
Quality standard92 percent calibration pass rate before an evaluator scores live

Why a generalist cannot benchmark a specialist model

Model evaluation looks like an annotation task and is priced like one, which is how most evaluation programmes go wrong. Ask a capable generalist to compare two answers about anticoagulant dosing, or two contract clauses, or two implementations of a caching layer, and they will give you a careful, considered, confident judgement based on the only thing they can actually assess. Which answer reads better.

Modern models are fluent by default. Both answers will be well structured, appropriately hedged and pleasant to read. The difference between them is whether the dosage is right, whether the clause survives contact with the relevant jurisdiction, whether the code has a race condition. None of that is visible without the underlying qualification.

A non-expert grading a specialist answer is not measuring correctness. They are measuring fluency, and reporting it as correctness.

This produces the worst kind of evaluation data, because it is confidently wrong rather than obviously missing. The scores look clean, agreement is often high, and the training signal steadily rewards the model for sounding more authoritative rather than for being more accurate.

Standing up expert panels at sprint speed

The operational problem in expert evaluation is not the evaluating. It is finding, verifying and contracting 180 qualified practitioners across eight fields, then doing it again when the client adds a ninth domain mid programme.

That is a recruitment problem, and it runs on the same infrastructure our staffing business uses: sourcing pipelines, credential verification, and contracting built for specialist hiring rather than for crowd sign up. It is why a new domain panel goes from brief to first scored batch in nine days. A vendor building this from an open contributor marketplace is starting that clock from zero every time.

Verification is the part that cannot be skipped. Domain evaluation attracts people who are adjacent to a field rather than practising in it, and adjacency is not detectable from the work itself, because an adjacent person also writes convincing justifications. It has to be caught at intake through credentials, and then confirmed through calibration.

Calibration before live scoring

No evaluator scores production data until they clear a calibration set of pre adjudicated items in their own domain, at a 92 percent pass rate.

The set is deliberately weighted toward items where the fluent answer and the correct answer diverge. Anyone selecting on polish fails it, which is exactly the failure mode the panel exists to prevent.

Making side by side comparison actually comparable

Comparing four to six models on the same prompt introduces distortions that have nothing to do with model quality, and every one of them has to be designed out before the numbers mean anything.

Model identity is hidden and output order is randomised per item, because evaluators develop expectations about a labelled model within a single session, and because position carries a measurable preference of its own. Length is controlled for in the rubric rather than left to instinct, since longer answers read as more thorough whether or not they are. Ties are permitted and tracked, because forcing a preference between two equivalent answers manufactures a signal that does not exist.

Each response is scored on separate dimensions rather than a single overall verdict: factual accuracy, domain appropriateness, completeness against the question asked, and safety and responsible AI criteria including refusal behaviour and bias. Separating them is what makes the output actionable. A single blended score tells a model builder that model C is behind. Dimension level scoring tells them model C is behind specifically on completeness in clinical prompts while leading on safety, which is a finding somebody can act on.

Disagreements between qualified experts are treated as findings rather than noise. When two credentialed practitioners split on the same item, that item is usually sitting on a genuine boundary of professional judgement, and those are the items a model builder most needs surfaced. They go to adjudication and then into the guideline set.

420K+ Expert evaluations delivered across eight complex domains.
6 Day sprint cadence, sustained alongside an active training cycle.
180 Credentialed practitioners scoring inside their own field.

Outcome and where the programme went next

The programme gave the model builder comparative, domain segmented benchmarks at a cadence fast enough to steer training rather than merely report on it. Because the evaluation ran every six days, a change made in one cycle could be tested in the next, and regressions in a specific domain surfaced while the cause was still identifiable.

Rubric design was tightened until the framework could reliably separate models differing by roughly three percentage points on a dimension. Anything coarser reads real improvement as noise, which is how evaluation programmes quietly stop influencing decisions.

The engagement then extended beyond evaluation into supervised fine tuning data and red teaming, both drawing on the same expert pool. That progression is the natural one. Once you have practitioners who can reliably identify what a model got wrong in their field, they are also the people best placed to write what it should have said, and to find the prompts where it fails.

Benchmarking models in domains where being wrong matters?

We build credentialed evaluator panels across 40+ domains and 50+ countries, with calibration, blind comparison design and adjudication built into the sprint.

Talk to our AI data team