Series B

A starting checklist for evaluation in regulated industries

What examiners actually look for, and the artefacts worth having before anyone asks.

22 August 2026Level 36 min read

Regulated deployment isn't a harder version of ordinary deployment. It's a different bar, and the difference is almost entirely about evidence that survives someone else reading it.

This isn't legal advice and requirements differ enormously by jurisdiction and sector. Treat it as a starting structure to take to the people who do know your regime.

The framing that matters

An examiner is not asking whether your model is good. They're asking whether your process was sound and whether you can prove it.

The five questions that keep coming up:

  1. How did you decide this system was fit for this particular use?
  2. Who signed off, and on what basis?
  3. What does it do when it's uncertain?
  4. How would you know if it degraded after deployment?
  5. Show me the record.

None is about model quality. All five are answerable with evaluation artefacts — or with narrative, which is what examiners are trained to discount.

Artefacts worth having before anyone asks

A written specification, dated and signed. What correct means for this workflow, in criteria a second person could apply. With the names of the people who approved it and when. Most organisations have never written this down anywhere, and reconstructing it later is transparently retrospective.

Evidence the graders agree with your experts. Sample size, agreement statistic, and the threshold you gated on. "We used an LLM to grade it" is not a methodology; "we calibrated against three senior practitioners on 200 held-out cases and gated at a 0.70 lower bound" is.

Named guardrails, scored as hard gates. The things that must never happen, with evidence they were tested rather than assumed. This is where most quality frameworks are silent and where examiners look first.

Documented escalation behaviour. What the system does when it's uncertain, and evidence it actually does that. Abstention rate and accuracy-conditional-on-answering are the two numbers worth having ready.

A change log with re-validation. Every model version, prompt change, and rubric revision, with the evaluation that followed it. A system that changed without re-validation is, from an examiner's perspective, an unvalidated system.

Provenance on the evaluation data. Where each case came from, who decided the correct answer, and what you're permitted to do with it.

Human oversight, evidenced. Not "a human reviews the output" but: which outputs, at what rate, by whom, with what override authority — and the override log showing it happens.

Two things that fail under examination

"We have logs." Logs can be edited, rotated, and selectively produced — none of which requires bad intent. What makes a record evidence is that it's append-only, attributable, and complete, meaning it contains the refusals and failures, not only the successes. A trail with no rejected specifications and no failed certifications in it is evidence that nothing was ever gated.

Self-attestation. If you built the benchmark, ran it, and reported the result, one party filled three roles. That's fine internally. It isn't evidence to someone external, however rigorous you actually were — because nobody outside can distinguish an honest self-assessment from a flattering one, and neither can you.

Sector-specific pressure points

Financial services. Model risk management frameworks generally expect independent validation, documented limitations, and ongoing monitoring. The phrase to be ready for is conceptual soundness — can you explain why this approach is appropriate for this use, not just that it scored well.

Healthcare administration. Distinguish administrative workflows from clinical decision support early and explicitly; the bar differs sharply and conflating them creates problems that are hard to unwind. For administrative work, payer-specific requirements and superseded criteria sets are the highest-yield failure modes and belong in the test set by name.

Insurance. Reasoning is the artefact under examination, not the decision. A correct outcome with an undocumented rationale is still an exposure. Weight your rubric accordingly.

Employment and credit decisions. Disparate impact analysis is typically a separate requirement from accuracy evaluation, with its own methodology. Don't assume a good pass rate addresses it — and involve counsel before designing the analysis, not after.

The starting sequence

If you're beginning from nothing, this order gets you defensible fastest:

  1. Write the specification and get it signed. Everything else references it.
  2. Name the guardrails — the must-nevers — and test them as hard gates.
  3. Build a small human-verified case set. Twenty is enough to start; it's the anchor for everything after.
  4. Calibrate your grader against those humans and record the number.
  5. Turn on the audit trail, including refusals, from day one. Retrofitting it is the expensive part.
  6. Schedule re-validation on a calendar, and on every version change.

Steps 1 through 4 are days, not months. Step 5 is the one that's cheap now and costly later, which is exactly why it gets deferred.

The one-line version

In a regulated setting the number is not the deliverable. The record of how you got the number is the deliverable — and it has to be legible to someone who has no reason to trust you and every reason to look carefully.