Conditionally
Approvedevidence pending · human review
Gradia Research
A visual walkthrough of the paper

Conditionally Approved

Proof-bound branchable universes for long-horizon AI agents under changing evidence, authority, and time.

Every AI agent benchmark you have seen scores a final answer. But real work happens in a world that changes while the agent acts — documents get retracted, authorities conflict, deadlines pass. This paper builds a testbed where every world change is cryptographically witnessed, runs four frontier models through it for $705.88, watches all 37 gradable attempts score exactly zero — and then, before blaming the models, shows that four AI judges given identical evidence can't even agree on whose fault that is.

Scroll the case file
01 · Exhibit A — The measurement problem

A plausible answer can still fail the work

Static benchmarks freeze the world, hand the agent a question, and grade the last thing it says. That works for trivia. It collapses for long-horizon agents, whose job unfolds over hundreds of actions while the ground truth shifts underneath them.

The paper opens with the failure modes a terminal score cannot see:

Invisible to a final-answer grader
Stale evidenceCites a source that was superseded mid-run.
Silent obedienceQuietly follows an unauthorized request.
Missed boundaryActs across a deadline it should have honored.
Epistemic residueKeeps using evidence that was retracted.
Toxic handoffHands off a summary that makes the next agent violate an active constraint.

And the converse holds too: a long, inefficient trajectory is not automatically wrong. The official score never penalizes step count — it asks only whether the submitted work is correct and fully evidence-grounded under a frozen contract.

The title carries a double meaning: in mortgage operations, approval is conditional on unresolved evidence. In this benchmark, every evaluation claim is likewise conditional — on a valid world, an admitted evaluator, and reviewable evidence.

02 · The apparatus — A universe you can prove

Every world change gets a witness

The primary object of study is not the mortgage tasks — it's the proof-bound, evolving, branchable Universe they run in. An episode isn't a transcript plus a score; it's a versioned scientific object, an eight-part identity frozen before anything is graded:

Ei = ( Ti, Mi, Si, Wi,0, Ai, Ωi, Ri, Ji )
Ti — frozen task edition Mi — requested & resolved model Si — scaffold & runtime identity Wi,0 — initial material world root Ai — ordered ledger of acts Ωi — occurrence-witness chain Ri — restore lineage Ji — evaluator edition & criteria

The crucial trick: the material world (what actually changed) and the agent-visible projection (what the agent could observe, and when) are separate cryptographic commitments. When an event fires at an eligible act boundary, the witness binds what changed, its before/after roots, who authorized it, and exactly which projection was delivered to the agent. Restores don't erase history — a restored world must name its parent and generation.

One witnessed episode — event, projection, restore lineage, judgment
act 1 read state act 2 calculate boundary b · event applied before/after roots witnessed projection delivered what the agent could see act 3 generation g restore → g+1 fork names its parent, generation & history terminal judge + receipt sibling branch

One event, one declared action boundary, one auditable evolution chain. "The agent was interrupted" becomes a set of falsifiable statements: the event was frozen, applied exactly once, the agent could observe the permitted projection, and the evaluator read the same edition. The witness proves what the agent could have seen — deliberately not what it attended to or believed.

The paper is careful about novelty: dynamic environments, snapshots, event logs, and provenance are all established prior art. The candidate contribution is the composition — binding world evolution, actor visibility, action boundaries, restore lineage, and admitted judgment tightly enough to catch defects that terminal state and ordinary logs miss. That incremental-value claim is stated as a hypothesis with a named falsifier, not a result.

03 · The panel — Five frozen task editions

Five trees, five ways the world moves

The testbed is a fully synthetic mortgage-underwriting workflow — no customer data, no real policy, and explicitly no claim of mortgage-domain validity. Mortgage is simply a demanding host: incomplete files, retractions, competing principals, hard deadlines, and handoffs. Each of the five frozen editions isolates one stressor:

Registered treatments · display names Alder – Birch
Alder
Document truth under pushback. Incomplete, conflicting document evidence.
Reconcile exact pages, versions, claims & current authority.
Cedar
Epistemic residue. Evidence observed, retracted, and replaced.
Stop using dead evidence; re-ground the decision.
Dogwood
Authority & fair judgment. Two legitimate principals conflict.
Apply authority scope, protect fairness, escalate correctly.
Fir
Temporal portfolio control. Coupled files evolve around a commitment boundary.
Coordinate scarce capacity, deadlines & post-change rechecks.
Birch
Honest handoff. Work crosses an agent/session boundary.
Quality scored partly by the successor's performance.

These are genuinely long-horizon: the safe scripted control policies need 113–150 actions under full load. That establishes solvability and a long execution path — deliberately not a claim of frontier difficulty.

04 · The contract — All gates or nothing

Reward is a conjunction, not a curve

Official reward is 1 only when every registered binary requirement passes, and 0 otherwise. Criterion vectors, branch probes, and narrative analyses are diagnostic evidence — never hidden partial credit. The product form means a strong final narrative cannot compensate for one stale citation or one unauthorized irreversible act. Try it:

Interactive · the strict binary conjunction
yi = k gi,k
y = 0 · FAIL

Click any gate to flip it. One zero anywhere zeroes the episode.

Before an episode may even enter that denominator, it needs a disposition: gradable, infrastructure censor, measurement defect, or pending adjudication. Only gradable attempts count against models — a transport failure can never masquerade as a model failure, and a retry is a new attempt, never an overwrite. The evaluator itself must survive mutation testing: positive controls, isolated one-defect probes, and post-change re-read gates.

05 · The run — 55 attempts, zero perfect

What actually happened when four frontier models tried

Four identity-pinned providers ran against the panel under a preregistered plan, later amended (openly, post-observation) to a hard cost-stop of 55 physical attempts. Every attempt kept its receipts:

All 55 physical attempts · disposition, then outcome
gradable (37) gradable · scored 0 (all 37) infrastructure exclusion (18)

The 18 exclusions: 7 OpenAI HTTP 401 auth refusals, 2 Anthropic HTTP 529 overloads, 6 read timeouts, 1 read error, 2 xAI responses with no admitted payload. They cost $173.35 and stay in the cost ledger — but never enter behavioral denominators. A naïve scoreboard would have reported them as model failures.

Execution inventory — explicitly not a ranking
ModelPhysicalGradableInfraPerfect passesSpend
Opus 514860$281.50
Gemini 3.1 Pro131300$2.99
GPT-5.6-sol181080$308.62
Grok 4.610640$112.77
Total5537180$705.88

Providers saw different cell mixes, the panel crossed three runtime cohorts, and the sample is small and outcome-incomplete — so the verifier refused to emit a balanced pass@2 statistic, any provider ranking, and every pass@5 claim.

Cell coverage · 20 provider × task cells
same-runtime pair (6) — pass@2 and pass² both false in all six one gradable attempt (9) — all zeros unobserved (5)

The cleanest slice: on Document Truth (Alder), all four providers supplied two gradable same-runtime attempts — eight attempts, zero passes. A complete result for one synthetic task; not extrapolable to the other four.

Identical zeros, wildly different trajectories · tool actions per selected pair
GPT-5.6 · Residue
303
GPT-5.6 · Document
200
Opus 5 · Document
150
Grok 4.6 · Document
145
Gemini 3.1 · Document
4
Gemini 3.1 · Residue
4

† Deterministic low-engagement marker: under 10% of the scripted control's actions with no source, attachment, or world event observed — 13/37 gradable traces overall, 7/21 selected primaries. In all ten of its Document and Residue attempts, Gemini queried branch.inspect and tried to submit after just two tool actions — a repeatable short-circuit through the branch surface. The paper explicitly leaves open whether that's model protocol failure, scaffold affordance ambiguity, or their interaction; the annotation changes no score and supports no efficiency or capability comparison.

A terminal scalar collapsed two orders of magnitude of execution difference into the same zero. The action ledger is what makes that difference reviewable.

06 · The stress test — Four judges, one file

Same evidence, four different verdicts

The 37 gradable episodes carried 1,375 machine-diagnostic criterion assignments: 486 green, 889 red. But a red surface is an unresolved observation, not yet a model-attributed failure — a non-engaging trace, an evaluator defect, or a disclosure defect can light up several criteria at once.

Could an LLM judge settle attribution? As a stress test, four independently pinned judges each classified all 889 red surfaces from identical frozen evidence as model failure, measurement defect, or unresolved. Every vote was advisory; zero scores changed.

Share of 889 surfaces labeled "model failure" — with each judge's mean stated confidence
Claude
Opus 5
28.3%
Grok 4.6
66.6%
GPT-5.6-sol
72.2%
Gemini 3.1
Pro Preview
83.9%
share labeled model failure mean stated confidence

Identical bytes, a 28.3%→83.9% spread — and confidence didn't help: the judge most eager to blame the model also reported the highest confidence (0.968). The result doesn't say who was right; it says one confident judge is not a stable attribution authority.

How often did the four judges agree? · 889 assignments
unanimous · 232 (26.1%) 3–1 majority · 384 2–1–1 plurality · 185 2–2 tie · 88
Descriptive Fleiss' κ = 0.151

It gets stranger: pairwise exact-label agreement ranged 35.4–69.5%, but the cited evidence anchors barely overlapped (Jaccard 0.089–0.292). Depending on the pair, 60.3–86.4% of exact-label agreements shared zero evidence anchors — consensus labels concealing different evidentiary paths. Proof-bound citations make that divergence inspectable; a scalar label cannot. The panel routed 682/889 surfaces (76.7%) to critical or high human-review priority: an accelerator, not an adjudicator.

Only 232 of 889 attributions were unanimous. The finding is not that plurality labels are truth — it's that disagreement itself became auditable.

07 · The ledger — What a valid attempt costs

The bill for 55 attempts

Long-horizon evaluation is physically expensive, and the paper treats cost as a first-class research question. Every call, token, and dollar — including those burned by infrastructure failures — stayed in the physical-attempt ledger:

Gradia Universes · physical-attempt ledger
frontier-v4 · 24 Aug 2026 · 14:24:51 UTC
provider calls0
input tokens0
output tokens0
official tool actions*0
official transcript turns*0
* from the 37 gradable episodes only
of which infra exclusions$173.35
provider spend$0.00

Scale is operational context, not a difficulty metric: a long trace can reveal expensive context accumulation or provider fragility, but only an admitted outcome establishes completion — and only a balanced, reviewed panel could support a broader comparison.

08 · The boundary — Claims and non-claims

What this paper does — and refuses — to claim

The paper's most distinctive feature may be its evidence discipline: overclaim refusal is executable, not editorial. The verifier itself refuses statistics the data can't support, and the manuscript keeps every conclusion on a claim ladder.

Established by this study

  • The instrument runs: 55 attempts frozen with recomputable identities, receipts, and dispositions.
  • 0/37 gradable attempts satisfied the exact perfect-rubric conjunction — an exact fact about these artifacts.
  • Infrastructure was material: 18/55 attempts were not valid model outcomes, and the censors caught them.
  • Identical zeros concealed ~100× differences in tool activity; the process ledger makes that reviewable.
  • One confident LLM judge is an unstable attribution authority (26.1% unanimity, κ = 0.151).
  • The verifier caught runtime drift and refused the pooled aggregate it would have contaminated.

Explicitly not claimed

  • A balanced model comparison, provider ranking, or reliability tier — different cell mixes, three runtime cohorts.
  • That models behaviorally failed: all 889 red surfaces await blinded human adjudication separating agent faults from measurement defects.
  • Frontier difficulty, tail reliability, or any pass@5 statement.
  • Real mortgage-domain validity — everything is synthetic.
  • Causal effects, training utility, or deception detection.
  • Novelty of hashing, logging, branching, or provenance in isolation — only the composition is the candidate contribution, with a named falsifier.
The claim ladder · evidence licenses only the lower rungs
1Implemented — canonical objects, replay, mutation probes, public index execute as describedreached
2Engineering-sensitive — planted safe/unsafe controls produce the intended deterministic distinctionsreached
3Observed preliminary machine results — 55 attempts, 37 gradable, exact denominators frozenreached
4Machine-attribution stability measured — four advisory judges; disagreement, not truthreached
5Human validity — two blinded reviewers × 889 red assignments, then adjudicationpending
6Calibrated difficulty, ranking, domain validity, causal attribution, noveltyunestablished

"No rhetorical section, including this one, may promote a result to a higher rung." The remaining gate for this frozen edition is narrow and concrete: two blinded human reviewers disposition every red criterion, disagreements go to a separately identified adjudicator, and agreement statistics are computed from sealed decisions.