GRADIA In plain English

For the person who signs off — not the person who codes

Before you let an AI
do a real job, read this.

Someone at your company wants to hand real work to an AI. You will be asked to approve it. This page explains — in plain words, no technical terms — how you know whether to say yes, and how you prove your answer later.

01 · The question

The question you're actually asking

It isn't “how smart is this AI?” It's simpler and harder than that:

“If I let an AI do this job, and it goes wrong, can I explain to my board what I knew and when I knew it?”

That is a question about proof, not about technology. Everything below is about getting you that proof.

MEMO · FOR DECISION Let the AI handle the loan write-ups? YES NO YOUR DESK
The decision lands on your desk either way. The only choice is what evidence sits under it.
02 · The score

The number you were shown doesn't answer it

The AI arrives with an impressive score. The score is real. But look at where it came from: someone else's test, built from someone else's tasks, graded by a judge nobody checked — on a day that has already passed.

You would never hire a person because they aced a different company's interview. This is the same thing. The number isn't wrong. It's just not about your job.

94% THEIR TEST THEIR TASKS · THEIR JUDGE YOUR JOB your files · your rules DOES 94% SURVIVE HERE?
A score earned somewhere else fades the moment it meets your actual work.
03 · The build

A test built from your own work — and a judge that passes its own exam first

Gradia builds the test out of your real work: your files, your systems, the standards your own people wrote down. Vague standards like “good tone” are turned away — they can't be checked. Then every task must prove it can be graded fairly before it counts. In one worked example, 4 of 12 tasks failed that proof. None of them shipped.

And the judge doing the grading? It is not trusted by default. It takes its own exam first: it may only grade your work after it agrees with your own experts often enough. In that same example, the first judge fell short — so it was not allowed to run. It was rebuilt until it agreed with the experts about 8 times in 10 on work it had never seen before.

Then it gets harder: people try to fool the judge on purpose, before it is trusted. When Gradia ran that trick-test against a popular public grading system, the grader was fooled 89 times in 1,800 tries — each one caught and recorded. Your judge gets the same treatment, and its result goes on your report.

YOUR REAL WORK files · standards YOUR TEST weak tasks refused THE JUDGE not trusted yet ITS OWN EXAM tricked on purpose only then may it grade
Worked example: agreement floor missed on the first try → judge rebuilt → ~8-in-10 agreement on unseen work. Measured: a public grading system fooled 89 / 1,800 tries.
04 · The proof

You walk away with a sealed result

When the test finishes, the result is sealed. Change even one line afterward and the seal visibly breaks — there is no way to quietly edit history. Your auditor, your regulator, your board: any of them can check the seal themselves, without calling Gradia. Nobody has to take our word for anything.

Try it yourself:

the result, as delivered
RESULT · SEALED passed 372 of 400 judge check 8 in 10 agree date locked in changes since none changes since ONE LINE EDITED INTACT BROKEN verifies rejected
Anyone with the file can check it. No phone call, no trust required — the seal either holds or it doesn't.
05 · The clock

Why this matters now: evidence for a real decision

When AI helps make a lending decision, reviewers need to understand the system, its limitations, and the evidence behind its use.

Colorado replaced its earlier AI framework with SB26-189, signed May 14, 2026. The enacted summary places developer technical-documentation duties at January 1, 2027 and describes record retention, notices and human review, subject to scope and exemptions. It does not support the annual testing mandate previously described here.

Gradia can support a buyer-approved assessment and review cadence. An evidence package does not by itself establish legal compliance. Colorado General Assembly, enacted summary · checked September 6, 2026.

TEST kept as proof BIG CHANGE rules or model TEST AGAIN kept as proof 2026 → ON
A proposed review cycle: reassess when the agreed workflow, policy or model changes.

Bring one sentence — the job you want done.
Leave with proof.

Your problem. Your tools. Your workflow. Tested against your own standards, graded by a judge your experts vouched for, sealed so anyone can check it later.

gradiahq.com