Gradia Research

Research that shows its work.

We study how to evaluate agents in worlds that change while they work, and how to tell a model failure from a broken ruler. Every public claim points to an exact edition, code artifact, evidence boundary, and result status.

01 · Published

Citable public records

Four DOI-backed research records, preserved with the code and study boundary needed to interpret them. Preliminary, pre-results, and one-seed labels stay visible because a public record should make unfinished inference easier to see, not easier to miss.

Paper 01

2 Sep 2026

Version 1.0.2

Published research artifact · one-seed paired GRPO

Reward Hacking in the RL Loop: Oracle-Witnessed Localization and Repair of Reward-Model Exploits During Training

A training-time study of what happens when an optimizer repeatedly probes a reward channel for exploitable seams. Matched 300-step GRPO/LoRA arms preserve exact manifests, frame chains, final adapters, analysis, and model-backed replay receipts so proxy success and oracle truth remain separately inspectable.

E1
Matched 300-step GRPO/LoRA arms · two verified 313-frame chains
E2
RLVR control gap 0 · gameable arm 58/64 proxy vs 1/64 oracle · gap 0.890625
E3
55 property/control checks · exact paired analysis · two model-backed final replays

Claim boundary: The frozen result is one seed on Qwen2.5-0.5B-Instruct and GSM8K. Both arms lost oracle accuracy, so the study establishes reward-channel exploitation and amplification in this cell—not capability improvement, population prevalence, a frontier-model result, or production-reward behavior. Repeated-seed variance and real-policy localization/repair remain open.

doi:10.5281/zenodo.22259605

Paper 02

1 Sep 2026

Published preprint · empirical evaluator study

The Reward-Hacking Wind Tunnel: An Evaluation Immune System that Attacks, Localizes, and Repairs Benchmark Scorers

An evaluation immune system that attacks the scorekeeper rather than the model. Oracle-wrong outputs probe frozen scorers, witnessed single-variable forks test causal attribution, guarded repairs are re-attacked on held-out evidence, and the resulting gameability and convergence measurements remain independently inspectable.

E1
Five public benchmark families · 962 L1 items · 357 oracle-witnessed exploits
E2
SWE-bench Verified · 100 live task containers · 1,800 grader-level attempts · 89 witnessed exploits
E3
165 property and control checks · offline instrument and report replay verified

Claim boundary: The paper reports scorer-gameability, causal-localization, repair-convergence, and live weak-grader results. It does not rank model capability or claim downstream training lift; human naturalness ratings and any optimizer-backed training study remain separate work.

doi:10.5281/zenodo.22233638

Paper 03

26 Aug 2026

Version 1.1.0

Preliminary results · human adjudication pending

Conditionally Approved: Proof-Bound Branchable Universes for Long-Horizon AI Agents Under Changing Evidence, Authority, and Time

A synthetic mortgage testbed for evaluating long-horizon agents while evidence, authority, and time change around them. The study binds each eligible episode to the exact world roots, visible projections, restore lineage, runtime, model, and evaluator that produced it.

E1
55 physical attempts · 37 gradable · 18 infrastructure exclusions
E2
1,375 machine-scored criterion surfaces
E3
Four-judge attribution stress test · Fleiss' κ 0.151

Claim boundary: The execution and four-judge stress test are complete. The paper does not claim a balanced provider ranking or model-attributed failure modes before blinded human adjudication.

doi:10.5281/zenodo.22104672

Paper 04

24 Aug 2026

Version 1.0.0

Pre-results · agreement study preregistered

The Value Engine Benchmark: An Evidence-Graded, Methodology-Controlled Environment for Multi-Touch Enterprise-Sales Negotiation

A fully synthetic enterprise-sales negotiation environment with a deterministic, evidence-graded harness and a frozen evaluation grid covering 13 models and 3,510 graded episodes.

E1
13-model frozen evaluation grid
E2
3,510 graded episodes in the declared grid
E3
Machine-generated statistics · checksums · CI drift gate

Claim boundary: The public edition establishes the instrument and frozen study design. Its preregistered judge-human agreement study has not yet run, so it makes no human-agreement claim.

doi:10.5281/zenodo.22073789

02 · Interactive

Explore the evidence systems

Six self-contained explainers make the machinery tangible. Change a branch, inspect a witness, replay a receipt, or follow a reward exploit from discovery to repair. The interaction helps explain the record; it never becomes the record.

Training experiment

Reward Hacking in the RL Loop

Walk through the paired GRPO arms, the proxy-oracle gap, and the evidence chain that keeps optimization from rewriting the result.

Boundary: Interactive interpretation of a one-seed study, not new evidence or a claim of capability improvement.

Open interactive explainer ↗

Evaluator metrology

The Reward-Hacking Wind Tunnel

Explore how oracle-wrong outputs attack a scorer, how witnessed forks localize the seam, and how a repair earns re-admission.

Boundary: Explains the published evaluator study; it does not add human naturalness or downstream training evidence.

Open interactive explainer ↗

Branchable benchmark

Conditionally Approved

Follow a long-horizon mortgage world as evidence, authority, deadlines, and agent-visible state change across witnessed branches.

Boundary: Preliminary machine results remain separate from model-attributed failure claims pending blinded human review.

Open interactive explainer ↗

Methods

Interruptible Universes

Inspect the evolution witness: the event, action boundary, world roots, projection, restore generation, and occurrence lineage.

Boundary: A methods explainer for a pre-results manuscript; confirmatory Study A remains unrun.

Open interactive explainer ↗

Long-horizon benchmark

The Value Engine Benchmark

See how a multi-touch sales negotiation is scored from evidence instead of sales-shaped prose, across evolving commitments and stakeholders.

Boundary: Explains the released instrument and frozen study design; the preregistered human-agreement study remains open.

Open interactive explainer ↗

Open verification layer

Gradia Guard

Replay a hash-chained receipt and see how model, tool, authorization, file, process, network, and side-effect evidence stays independently verifiable.

Boundary: A product explainer for the open verifier, not a certificate, managed-service attestation, or substitute for a deployment receipt.

Open interactive explainer ↗

03 · Methods

Pre-results methods manuscript

Interruptible Universes: Evolution Witnesses for Verifiable World Change in Agent Benchmarks

Introduces an evolution witness that joins the declared event, action boundary, before-and-after world roots, exact agent-visible projection, restore generation, and prior occurrence head. The public artifact validates deterministic controls and verifier mutations; confirmatory Study A remains unrun.

04 · Program

The questions behind the papers

  1. RQ01

    World change

    Can a receipt prove what changed, when it became visible, and whether snapshot or restore preserved the same history?

  2. RQ02

    Counterfactual branches

    Can paired worlds isolate the effect of retracted evidence, authority, timing, or observation without changing unrelated state?

  3. RQ03

    Evaluator validity

    Which verdicts survive grader mutations, evidence-view ablations, and meaning-preserving changes in presentation?

  4. RQ04

    Human calibration

    Where do qualified reviewers agree, where do model judges disagree, and which surfaces must remain unresolved?

  5. RQ05

    Evaluator autoimmunity

    Can an evaluator reject shortcut patches that improve its attack score by damaging honest controls, naturalness, or held-out validity?

  6. RQ06

    Training steering

    Can verified failure evidence become a rights-safe, split-clean, controlled training study without mistaking a proposed intervention for measured model lift?

These are research questions, not completed findings. The published paper and each manuscript state which instrument tests are complete, which empirical work remains, and which claims the evidence cannot support.

Explore released Universes →