Proof-bound branchable universes for long-horizon AI agents under changing evidence, authority, and time.
Every AI agent benchmark you have seen scores a final answer. But real work happens in a world that changes while the agent acts — documents get retracted, authorities conflict, deadlines pass. This paper builds a testbed where every world change is cryptographically witnessed, runs four frontier models through it for $705.88, watches all 37 gradable attempts score exactly zero — and then, before blaming the models, shows that four AI judges given identical evidence can't even agree on whose fault that is.
Static benchmarks freeze the world, hand the agent a question, and grade the last thing it says. That works for trivia. It collapses for long-horizon agents, whose job unfolds over hundreds of actions while the ground truth shifts underneath them.
The paper opens with the failure modes a terminal score cannot see:
And the converse holds too: a long, inefficient trajectory is not automatically wrong. The official score never penalizes step count — it asks only whether the submitted work is correct and fully evidence-grounded under a frozen contract.
The title carries a double meaning: in mortgage operations, approval is conditional on unresolved evidence. In this benchmark, every evaluation claim is likewise conditional — on a valid world, an admitted evaluator, and reviewable evidence.
The primary object of study is not the mortgage tasks — it's the proof-bound, evolving, branchable Universe they run in. An episode isn't a transcript plus a score; it's a versioned scientific object, an eight-part identity frozen before anything is graded:
The crucial trick: the material world (what actually changed) and the agent-visible projection (what the agent could observe, and when) are separate cryptographic commitments. When an event fires at an eligible act boundary, the witness binds what changed, its before/after roots, who authorized it, and exactly which projection was delivered to the agent. Restores don't erase history — a restored world must name its parent and generation.
One event, one declared action boundary, one auditable evolution chain. "The agent was interrupted" becomes a set of falsifiable statements: the event was frozen, applied exactly once, the agent could observe the permitted projection, and the evaluator read the same edition. The witness proves what the agent could have seen — deliberately not what it attended to or believed.
The paper is careful about novelty: dynamic environments, snapshots, event logs, and provenance are all established prior art. The candidate contribution is the composition — binding world evolution, actor visibility, action boundaries, restore lineage, and admitted judgment tightly enough to catch defects that terminal state and ordinary logs miss. That incremental-value claim is stated as a hypothesis with a named falsifier, not a result.
The testbed is a fully synthetic mortgage-underwriting workflow — no customer data, no real policy, and explicitly no claim of mortgage-domain validity. Mortgage is simply a demanding host: incomplete files, retractions, competing principals, hard deadlines, and handoffs. Each of the five frozen editions isolates one stressor:
These are genuinely long-horizon: the safe scripted control policies need 113–150 actions under full load. That establishes solvability and a long execution path — deliberately not a claim of frontier difficulty.
Official reward is 1 only when every registered binary requirement passes, and 0 otherwise. Criterion vectors, branch probes, and narrative analyses are diagnostic evidence — never hidden partial credit. The product form means a strong final narrative cannot compensate for one stale citation or one unauthorized irreversible act. Try it:
Click any gate to flip it. One zero anywhere zeroes the episode.
Before an episode may even enter that denominator, it needs a disposition: gradable, infrastructure censor, measurement defect, or pending adjudication. Only gradable attempts count against models — a transport failure can never masquerade as a model failure, and a retry is a new attempt, never an overwrite. The evaluator itself must survive mutation testing: positive controls, isolated one-defect probes, and post-change re-read gates.
Four identity-pinned providers ran against the panel under a preregistered plan, later amended (openly, post-observation) to a hard cost-stop of 55 physical attempts. Every attempt kept its receipts:
The 18 exclusions: 7 OpenAI HTTP 401 auth refusals, 2 Anthropic HTTP 529 overloads, 6 read timeouts, 1 read error, 2 xAI responses with no admitted payload. They cost $173.35 and stay in the cost ledger — but never enter behavioral denominators. A naïve scoreboard would have reported them as model failures.
| Model | Physical | Gradable | Infra | Perfect passes | Spend |
|---|---|---|---|---|---|
| Opus 5 | 14 | 8 | 6 | 0 | $281.50 |
| Gemini 3.1 Pro | 13 | 13 | 0 | 0 | $2.99 |
| GPT-5.6-sol | 18 | 10 | 8 | 0 | $308.62 |
| Grok 4.6 | 10 | 6 | 4 | 0 | $112.77 |
| Total | 55 | 37 | 18 | 0 | $705.88 |
Providers saw different cell mixes, the panel crossed three runtime cohorts, and the sample is small and outcome-incomplete — so the verifier refused to emit a balanced pass@2 statistic, any provider ranking, and every pass@5 claim.
The cleanest slice: on Document Truth (Alder), all four providers supplied two gradable same-runtime attempts — eight attempts, zero passes. A complete result for one synthetic task; not extrapolable to the other four.
† Deterministic low-engagement marker: under 10% of the scripted control's actions with no source, attachment, or world event observed — 13/37 gradable traces overall, 7/21 selected primaries. In all ten of its Document and Residue attempts, Gemini queried branch.inspect and tried to submit after just two tool actions — a repeatable short-circuit through the branch surface. The paper explicitly leaves open whether that's model protocol failure, scaffold affordance ambiguity, or their interaction; the annotation changes no score and supports no efficiency or capability comparison.
A terminal scalar collapsed two orders of magnitude of execution difference into the same zero. The action ledger is what makes that difference reviewable.
The 37 gradable episodes carried 1,375 machine-diagnostic criterion assignments: 486 green, 889 red. But a red surface is an unresolved observation, not yet a model-attributed failure — a non-engaging trace, an evaluator defect, or a disclosure defect can light up several criteria at once.
Could an LLM judge settle attribution? As a stress test, four independently pinned judges each classified all 889 red surfaces from identical frozen evidence as model failure, measurement defect, or unresolved. Every vote was advisory; zero scores changed.
Identical bytes, a 28.3%→83.9% spread — and confidence didn't help: the judge most eager to blame the model also reported the highest confidence (0.968). The result doesn't say who was right; it says one confident judge is not a stable attribution authority.
It gets stranger: pairwise exact-label agreement ranged 35.4–69.5%, but the cited evidence anchors barely overlapped (Jaccard 0.089–0.292). Depending on the pair, 60.3–86.4% of exact-label agreements shared zero evidence anchors — consensus labels concealing different evidentiary paths. Proof-bound citations make that divergence inspectable; a scalar label cannot. The panel routed 682/889 surfaces (76.7%) to critical or high human-review priority: an accelerator, not an adjudicator.
Only 232 of 889 attributions were unanimous. The finding is not that plurality labels are truth — it's that disagreement itself became auditable.
Long-horizon evaluation is physically expensive, and the paper treats cost as a first-class research question. Every call, token, and dollar — including those burned by infrastructure failures — stayed in the physical-attempt ledger:
Scale is operational context, not a difficulty metric: a long trace can reveal expensive context accumulation or provider fragility, but only an admitted outcome establishes completion — and only a balanced, reviewed panel could support a broader comparison.
The paper's most distinctive feature may be its evidence discipline: overclaim refusal is executable, not editorial. The verifier itself refuses statistics the data can't support, and the manuscript keeps every conclusion on a claim ladder.
"No rhetorical section, including this one, may promote a result to a higher rung." The remaining gate for this frozen edition is narrow and concrete: two blinded human reviewers disposition every red criterion, disagreements go to a separately identified adjudicator, and agreement statistics are computed from sealed decisions.