Not on a public benchmark. On your workflow, against your policy, using your systems — with a result you can hand to your board, your regulator, or someone who wasn’t in the room.
94.2
any public score · illustrative
measured on someone else’s tasks
by someone else’s grader
on a day that has already passed
The number you were shown is real. It is just not about your job.
worked example: end-to-end mortgage underwriting · specimen data from a dogfood run unless a section is marked MEASURED
01 · Your ruler
The first thing built is your ruler, not a model.
✓criteria your underwriters wrote, made checkable by a machine
✓your 43% DTI overlay, your exception classes, your policy manual
✗“reasonably clear” and “good tone” — refused at the gate. Not criteria.
Frozen
spec v0.3 sha256:7c1e9a04d2…f31b
immutable · every later artifact inherits this fingerprint
a superseded spec is a different benchmark — on purpose
Then a world that mirrors your systems — isolated, pinned, no egress; nothing touches production. Any run can be paused, branched, and re-entered, byte for byte. A transcript tells you what happened. A world you can re-enter tells you why.
every event from the first sentence onward lands on a hash-chained audit trail · expert review and an AI quality gate run before anything freezes
02 · The judge, attacked first
The AI judge is not trusted by default.
agreement with your annotators · as first built0.61
✗ wilson 95% lower bound 0.61 — under the 0.70 floor. It does not run.
rewritten against your experts’ reasoning · unseen holdout0.805
✓ clears the floor — and agreement is necessary, not sufficient. A judge can agree and still be gameable.
So the judge is attacked on purpose, before it is trusted. The same instrument was pointed at a live public grader:
Measured · Wind Tunnel
public grader, live, as shipped49.4exploits / 1,000 · 95% [36.7, 62.8]
GSM8K control0.0 at every budget — the control the instrument must pass
1,800 attempts. 89 witnessed exploits — answers the grader marked correct that an oracle proves wrong. And because patches usually move the hole rather than close it, every repair is re-measured too: the certificate carries the number, not the intent.
the same tunnel runs on your judge · its own gameability goes on your certificate · a judge that drifts below its floor is deactivated
03 · One ruler, honestly applied
Then the models compete — and every run leaves a receipt.
Anthropic, OpenAI, Google, xAI, open weights, your current stack, your fine-tune. Same 400 certified tasks. Same world. Same judge. Same budget cap.
Specimen data
claude-opus-50.91[0.88, 0.94]
gpt-5.6-sol0.87[0.83, 0.90]
gemini-3.1-pro0.83[0.79, 0.87]
your current stack0.76[0.71, 0.81]
Confidence intervals on every bar — so a gap that is not real cannot look real.
specimen data — your run reports your models · per-episode model pins, so a mid-run provider update cannot change the ruler · a $0.00 all-zero run is a broken gateway, not a result, and the receipt says which
04 · Before you spend
Before you buy data, prove the failure is real.
The largest failure cluster — self-employed income averaged from one year, 61 episodes — is re-entered in the recorded world and forked once. One variable changes. Everything else holds still.
61 failures · re-entered · same seed
Witnessed
two-year transcript added → 58 of 61 pass
Control
world unchanged → 0 of 61 pass
The missing document is the cause — and the data ask is sized from a witnessed count, not a hunch.
One ruler then answers four questions: which model to use, where it still fails, what data would fix it, and — after the fix — the identical fingerprint is re-run, so the delta you see is the model, not the test.
every failure cluster is forked once before it becomes a data ask · a fork proves a cause, not a correlation
05 · The certificate
You leave with proof.
Mortgage Underwriting Agent Benchmark
headline pass rate93%
judge calibrationwilson lb 0.805 · n 200
judge gameability4.1 / 1,000 · re-measured after repair
spec + world + bundle rootsha256:7c1e9a04d2…
statussigned · anchored · verifies offline
Change one digit and the signature stops verifying. It checks offline, with no call home — your auditor can verify it without ever talking to us.
no override flag, no manual issuance, ever · anyone with the evidence bundle can recompute it
Your problem. Your tools. Your workflow. Measured.
A frozen spec your experts signed off on. A world that mirrors your systems. Four hundred certified tasks. A judge that agrees with your people and was attacked before it was trusted. A signed answer to the only question that ever mattered: can an agent do this job, here, for you.