Explainers

Understand the platform, then inspect the research behind it. Start with the enterprise journey or the detailed technical reference. Research pages retain their own methods, evidence and claim boundaries.

Start here

Guided walkthrough

How Gradia works

From an existing agent, observed human work or a structured brief to approved rules, inspectable attempts and an evidence-based decision.

Explore →
Technical reference

How Gradia works, in depth

Capture boundaries, versioned history, executable worlds, Spec → IR → Build, governed calls and the native evidence lifecycle. Diagrams, contracts and source references.

Explore →
Plain English

For the person who signs off

Zero jargon: the board question, why borrowed scores don't answer it, and the sealed result your auditor can check.

Explore →
For buyers

Gradia for executives

Six beats, ninety seconds of scroll: your ruler, the judge attacked first, and proof that verifies offline.

Explore →

The research program

Pillar 3 · measured

The Reward-Hacking Wind Tunnel

Attack the scorer, localize the flaw, measure whether repair cures or relocates. 89 witnessed exploits on a live public grader.

Explore →
Pillar 4 · measured

Reward Hacking in the RL Loop

The Goodhart gap opening live in training — proxy 0.98, true quality 0.00 — and the detector that fires before saturation.

Explore →
Pillar 2

Gradia Guard

Hash-chained evidence you can try to break — and exactly what a green verify does and does not prove.

Explore →
Pre-results draft

Interruptible Universes

Evolution witnesses for verifiable world change — with its preregistered studies and empty result tables published in advance.

Explore →
Under review

Conditionally Approved

Proof-bound branchable universes for long-horizon agents — 55 attempts, every zero audited, judges that couldn't agree.

Explore →
Benchmark

The Value Engine Benchmark

Evidence-graded enterprise-sales negotiation: gated facts, cite-or-abstain judging, matched-pair methodology lift.

Explore →