Gradia logo Gradia logo GRADIA

Gradia · for the decision

Can an agent actually do the job you’d hire for?

Not on a public benchmark. On your workflow, against your policy, using your systems — with a result you can hand to your board, your regulator, or someone who wasn’t in the room.

94.2
any public score · illustrative
measured on someone else’s tasks
by someone else’s grader
on a day that has already passed

The number you were shown is real. It is just not about your job.

worked example: end-to-end mortgage underwriting · specimen data from a dogfood run unless a section is marked MEASURED

01 · Your ruler

The first thing built is your ruler, not a model.

Frozen
spec v0.3    sha256:7c1e9a04d2…f31b
immutable · every later artifact inherits this fingerprint
a superseded spec is a different benchmark — on purpose

Then a world that mirrors your systems — isolated, pinned, no egress; nothing touches production. Any run can be paused, branched, and re-entered, byte for byte. A transcript tells you what happened. A world you can re-enter tells you why.

every event from the first sentence onward lands on a hash-chained audit trail · expert review and an AI quality gate run before anything freezes

02 · The judge, attacked first

The AI judge is not trusted by default.

agreement with your annotators · as first built0.61
✗  wilson 95% lower bound 0.61 — under the 0.70 floor. It does not run.
rewritten against your experts’ reasoning · unseen holdout0.805
✓  clears the floor — and agreement is necessary, not sufficient. A judge can agree and still be gameable.

So the judge is attacked on purpose, before it is trusted. The same instrument was pointed at a live public grader:

Measured · Wind Tunnel
public grader, live, as shipped49.4 exploits / 1,000 · 95% [36.7, 62.8]
GSM8K control0.0 at every budget — the control the instrument must pass

1,800 attempts. 89 witnessed exploits — answers the grader marked correct that an oracle proves wrong. And because patches usually move the hole rather than close it, every repair is re-measured too: the certificate carries the number, not the intent.

the same tunnel runs on your judge · its own gameability goes on your certificate · a judge that drifts below its floor is deactivated

03 · One ruler, honestly applied

Then the models compete — and every run leaves a receipt.

Anthropic, OpenAI, Google, xAI, open weights, your current stack, your fine-tune. Same 400 certified tasks. Same world. Same judge. Same budget cap.

Specimen data
claude-opus-50.91 [0.88, 0.94]
gpt-5.6-sol0.87 [0.83, 0.90]
gemini-3.1-pro0.83 [0.79, 0.87]
your current stack0.76 [0.71, 0.81]

Confidence intervals on every bar — so a gap that is not real cannot look real.

specimen data — your run reports your models · per-episode model pins, so a mid-run provider update cannot change the ruler · a $0.00 all-zero run is a broken gateway, not a result, and the receipt says which

04 · Before you spend

Before you buy data, prove the failure is real.

The largest failure cluster — self-employed income averaged from one year, 61 episodes — is re-entered in the recorded world and forked once. One variable changes. Everything else holds still.

61 failures · re-entered · same seed
Witnessed
two-year transcript added
58 of 61 pass
Control
world unchanged
0 of 61 pass

The missing document is the cause — and the data ask is sized from a witnessed count, not a hunch.

One ruler then answers four questions: which model to use, where it still fails, what data would fix it, and — after the fix — the identical fingerprint is re-run, so the delta you see is the model, not the test.

every failure cluster is forked once before it becomes a data ask · a fork proves a cause, not a correlation

05 · The certificate

You leave with proof.

Mortgage Underwriting Agent Benchmark
headline pass rate93%
judge calibrationwilson lb 0.805 · n 200
judge gameability4.1 / 1,000 · re-measured after repair
spec + world + bundle rootsha256:7c1e9a04d2…
statussigned · anchored · verifies offline

Change one digit and the signature stops verifying. It checks offline, with no call home — your auditor can verify it without ever talking to us.

no override flag, no manual issuance, ever · anyone with the evidence bundle can recompute it

Your problem. Your tools.
Your workflow. Measured.

A frozen spec your experts signed off on. A world that mirrors your systems. Four hundred certified tasks. A judge that agrees with your people and was attacked before it was trusted. A signed answer to the only question that ever mattered: can an agent do this job, here, for you.

gradiahq.com
Gradiabring one sentence · leave with a certificate