Gradia Evals · Public historical runs

Find the agent you can trust.

Explore real published runs. Compare outcomes, inspect failures, and trace the gap back to the work.

Published benchmark runs · τ2-bench · recorded cost and performance

400 runs. 50 tasks. Every outcome visible.GPT-5.2 · high and Claude Opus 4.5 · high · source-reported rewards

Same task specs, trial keys, recorded seeds and shared recorded configuration. Historical runtime dependencies were not independently reproduced.

GPT-5.2 · high

83.0%recorded pass · 166 / 200 runs

Claude Opus 4.5 · high

84.0%recorded pass · 168 / 200 runs

Task-level disagreement

14 / 50tasks with different observed pass rates

Variable across trials

17 / 50tasks where at least one model both passes and fails

01 · Outcome comparison

How often did it work?

Source rewards
GPT-5.2 · high166 pass · 34 fail · 0 unknown
83.0%
Claude Opus 4.5 · high168 pass · 32 fail · 0 unknown
84.0%
0%25%50%75%100%

Pass = upstream reward 1. All selected runs stay in the denominator. Published runs are not a current frontier ranking or a forecast of your workflow.

02 · Matched outcomes

Where do they disagree?

200 pairs
155Both pass
11Only GPT-5.2 · high passes
13Only Claude Opus 4.5 · high passes
21Both fail

Shared task/trial keys only, even in the full-file view. 0 pairs have unknown outcomes. Disagreement identifies a next test; it is not causal attribution.

Measured trade-offs · source records

What does each successful outcome cost?

Mean recorded agent cost versus pass rate. Every selected run, including failures, contributes to spend.

Recorded model-priced costs
0%25%50%75%100%$0.00$0.12$0.24$0.36$0.48Mean recorded agent USD / runGPT-5.2 · highClaude Opus 4.5 · high

Lower cost and higher pass rate describe a trade-off within this source cohort. The chart does not establish statistical significance or invoice-verified billing.

Selected model

GPT-5.2 · high

Recorded mean agent cost
$0.1138
Observed spend / recorded pass
$0.1371
Recorded pass rate
83.0% · 166/200
Cost coverage
200/200 runs
Simulation duration · median / p95
870.9s / 1,443.1s

Duration includes the simulated user and tool work, not just agent inference. Cost excludes the user simulator. Observed spend per pass includes spending on failed runs; it is not a forecast.

Exact cost and outcome comparison
ModelPassed / selectedMean agent costSpend / passMedian / p95 simulation duration
166/200$0.1138 · 200/200$0.1371870.9s / 1,443.1s · 200/200
168/200$0.3992 · 200/200$0.4752128.5s / 240.4s · 200/200

03 · Task-level resolution

Averages hide the tasks that break.

Each column is one task. Select a cell to inspect its runs below.

Task423545781415192124293133370123456910111213161718202223252627283032343638394041434446474849
GPT-5.2 · high
Claude Opus 4.5 · high
Gap · pp+75-50+50+25-25-25-25+25-25+25+25-25-25+25000000000000000000000000000000000000
GPT-5.2 · highClaude Opus 4.5 · highStronger model color = more passes · gap = second model minus first

Task pass-rate distribution

Task difficulty profiles

One task, one band per model. See where reliable passes remain rare in these recorded trials.

0–25%>25–40%>40–60%>60%
GPT-5.2 · high50/50 tasks classified
GPT-5.2 · high: 7 of 50 tasks in 0–25%. At most one in four retained trials passed.GPT-5.2 · high: 1 of 50 tasks in >40–60%. More than 40%, at most 60% of retained trials passed.GPT-5.2 · high: 42 of 50 tasks in >60%. More than 60% of retained trials passed.
Claude Opus 4.5 · high50/50 tasks classified
Claude Opus 4.5 · high: 6 of 50 tasks in 0–25%. At most one in four retained trials passed.Claude Opus 4.5 · high: 2 of 50 tasks in >40–60%. More than 40%, at most 60% of retained trials passed.Claude Opus 4.5 · high: 42 of 50 tasks in >60%. More than 60% of retained trials passed.
Point to or focus a band to inspect its task count. Select it to see the exact tasks.

Bands use each model’s pass count divided by all retained trials for a task. Unknown outcomes or missing model/task runs stay unclassified. This describes observed results, not intrinsic task difficulty.

Exact difficulty counts and task shares
Common population: 50 tasks per model
Model / bandTasksShare of tasks
GPT-5.2 · high0–25%7 / 507/50 (14.0%)
GPT-5.2 · high>25–40%0 / 500/50 (0.0%)
GPT-5.2 · high>40–60%1 / 501/50 (2.0%)
GPT-5.2 · high>60%42 / 5042/50 (84.0%)
GPT-5.2 · highIncomplete outcomes0 / 500/50 (0.0%)
Claude Opus 4.5 · high0–25%6 / 506/50 (12.0%)
Claude Opus 4.5 · high>25–40%0 / 500/50 (0.0%)
Claude Opus 4.5 · high>40–60%2 / 502/50 (4.0%)
Claude Opus 4.5 · high>60%42 / 5042/50 (84.0%)
Claude Opus 4.5 · highIncomplete outcomes0 / 500/50 (0.0%)

Percent labels are rounded to one decimal; counts give the exact shares. Four bands: 0–25%, >25–40%, >40–60%, >60%.

Different work, different outcomes

Workflow categories

Compare both models on the same task set in each category. Every task belongs to exactly one category.

GPT-5.2 · highClaude Opus 4.5 · high
Recorded pass rate
0%50%100%
GPT-5.2 · high, Cancellation: 8 of 28 retained runs pass; 0 unknown outcomes; 7 of 7 tasks covered.Claude Opus 4.5 · high, Cancellation: 13 of 28 retained runs pass; 0 unknown outcomes; 7 of 7 tasks covered.
GPT-5.2 · high, Booking: 19 of 20 retained runs pass; 0 unknown outcomes; 5 of 5 tasks covered.Claude Opus 4.5 · high, Booking: 17 of 20 retained runs pass; 0 unknown outcomes; 5 of 5 tasks covered.
GPT-5.2 · high, Reservation changes: 47 of 56 retained runs pass; 0 unknown outcomes; 14 of 14 tasks covered.Claude Opus 4.5 · high, Reservation changes: 45 of 56 retained runs pass; 0 unknown outcomes; 14 of 14 tasks covered.
GPT-5.2 · high, Read-only / other: 92 of 96 retained runs pass; 0 unknown outcomes; 24 of 24 tasks covered.Claude Opus 4.5 · high, Read-only / other: 93 of 96 retained runs pass; 0 unknown outcomes; 24 of 24 tasks covered.
Focus a marker for its exact pass count and run denominator. Select a category to see its tasks.

All tasks here are airline tasks. Categories are editorial labels from the task’s reference action tools, not independent benchmark domains or the agent’s observed actions.

Exact category rates, gaps and coverage
Gap = Claude Opus 4.5 · high minus GPT-5.2 · high, percentage points
CategoryGPT-5.2 · highClaude Opus 4.5 · highUnrounded gap
Cancellation7 tasks8/28 pass0 unknown · 7/7 tasks13/28 pass0 unknown · 7/7 tasks+17.857142857142858 pp
Booking5 tasks19/20 pass0 unknown · 5/5 tasks17/20 pass0 unknown · 5/5 tasks-10 pp
Reservation changes14 tasks47/56 pass0 unknown · 14/14 tasks45/56 pass0 unknown · 14/14 tasks-3.5714285714285716 pp
Read-only / other24 tasks92/96 pass0 unknown · 24/24 tasks93/96 pass0 unknown · 24/24 tasks+1.0416666666666667 pp

Unknown outcomes remain in run denominators. A gap is unavailable if either model lacks runs for a task in that category. Trial counts may differ; these are descriptive rates without significance claims.

How task profiles and workflow categories are defined

Each model uses the same 50-task population. Difficulty bands count tasks equally; category rates count retained runs. The selected population contains 400 runs. Switching between matched and complete files can change trial counts and the resulting rates.

Category priority is fixed and mutually exclusive. A task with several reference actions takes the first matching rule below. If reference tool sets disagree across records, it is shown as unclassified; the model’s actual tool calls never determine its category.

  1. Cancellation. A reference action calls cancel_reservation.
  2. Booking. Otherwise, a reference action calls book_reservation.
  3. Reservation changes. Otherwise, a reference action calls a tool beginning update_reservation_.
  4. Refund / compensation. Otherwise, a reference action calls send_certificate, refund_reservation, or a tool beginning refund_.
  5. Read-only / other. All remaining reference action sets, including empty sets.
  6. Unclassified reference. Reference tool sets disagree between retained records for this task.

Method: gradia.public-reference-workflow-categories.v1. Classification does not establish policy compliance or task success. Identical task specifications do not establish identical user simulation, seeds or harness versions.

Shared failure pocket

3 tasks never pass for either model.

Every selected trial fails for both. Start with the task requirements and the tool sequence before changing the model.

Repeat-call signal

39 runs repeat an identical tool call.

Observed in 400/400 runs. Repetition can be a valid retry; inspect the sequence before treating it as waste.

Decision boundary

A benchmark pass is one piece of evidence.

These outcomes use the publisher’s reward. Your approval rules, business costs and recovery requirements need their own checks.

See which evidence checks catch a gap →

04 · Effort and outcomes

Longer isn’t automatically better.

PassFailAssistant-message count (not elapsed time)

Buckets mix both models. Task difficulty and run length can be related; this is descriptive, not evidence that extra steps cause failure.

05 · Tool footprint

Which tools shape the work?

get reservation details
95.0%
95.5%
get user details
81.0%
93.0%
search direct flight
50.0%
44.5%
transfer to human agents
46.5%
31.5%
get flight status
41.0%
15.5%
search onestop flight
29.5%
26.0%
update reservation flights
24.5%
25.5%
book reservation
14.5%
14.5%

Top 8 tools by run coverage; each pair follows the model order above. One run may call several tools. Names alone do not establish correct arguments or results.

06 · Follow a point into the work

Inspect the runs behind the chart.

Filter, select a point, then follow its recorded tool sequence.

400 runs selected · 400 with both chart measurements

006310126201883025140Assistant messagesTool callsGPT-5.2 · high · task 46 · trial 1 · passGPT-5.2 · high · task 46 · trial 0 · passGPT-5.2 · high · task 36 · trial 0 · passGPT-5.2 · high · task 43 · trial 1 · passGPT-5.2 · high · task 13 · trial 0 · passGPT-5.2 · high · task 13 · trial 1 · passGPT-5.2 · high · task 43 · trial 0 · passGPT-5.2 · high · task 28 · trial 0 · passGPT-5.2 · high · task 27 · trial 1 · passGPT-5.2 · high · task 0 · trial 1 · passGPT-5.2 · high · task 9 · trial 1 · passGPT-5.2 · high · task 9 · trial 0 · passGPT-5.2 · high · task 36 · trial 1 · passGPT-5.2 · high · task 40 · trial 0 · passGPT-5.2 · high · task 41 · trial 0 · passGPT-5.2 · high · task 27 · trial 0 · passGPT-5.2 · high · task 23 · trial 0 · failGPT-5.2 · high · task 1 · trial 0 · passGPT-5.2 · high · task 40 · trial 1 · passGPT-5.2 · high · task 3 · trial 0 · passGPT-5.2 · high · task 48 · trial 0 · passGPT-5.2 · high · task 1 · trial 1 · passGPT-5.2 · high · task 5 · trial 1 · passGPT-5.2 · high · task 18 · trial 0 · passGPT-5.2 · high · task 41 · trial 1 · passGPT-5.2 · high · task 49 · trial 1 · passGPT-5.2 · high · task 29 · trial 0 · failGPT-5.2 · high · task 34 · trial 0 · passGPT-5.2 · high · task 26 · trial 1 · passGPT-5.2 · high · task 49 · trial 0 · passGPT-5.2 · high · task 28 · trial 1 · passGPT-5.2 · high · task 26 · trial 0 · passGPT-5.2 · high · task 3 · trial 1 · passGPT-5.2 · high · task 25 · trial 0 · passGPT-5.2 · high · task 11 · trial 1 · passGPT-5.2 · high · task 12 · trial 1 · passGPT-5.2 · high · task 16 · trial 0 · passGPT-5.2 · high · task 5 · trial 0 · passGPT-5.2 · high · task 25 · trial 1 · passGPT-5.2 · high · task 23 · trial 1 · failGPT-5.2 · high · task 17 · trial 0 · passGPT-5.2 · high · task 24 · trial 0 · passGPT-5.2 · high · task 37 · trial 1 · passGPT-5.2 · high · task 30 · trial 0 · passGPT-5.2 · high · task 45 · trial 0 · failGPT-5.2 · high · task 15 · trial 0 · passGPT-5.2 · high · task 21 · trial 0 · passGPT-5.2 · high · task 20 · trial 1 · passGPT-5.2 · high · task 8 · trial 0 · passGPT-5.2 · high · task 17 · trial 1 · passGPT-5.2 · high · task 0 · trial 0 · passGPT-5.2 · high · task 4 · trial 1 · passGPT-5.2 · high · task 12 · trial 0 · passGPT-5.2 · high · task 48 · trial 1 · passGPT-5.2 · high · task 30 · trial 1 · passGPT-5.2 · high · task 37 · trial 0 · passGPT-5.2 · high · task 8 · trial 1 · passGPT-5.2 · high · task 22 · trial 0 · passGPT-5.2 · high · task 29 · trial 1 · failGPT-5.2 · high · task 2 · trial 1 · passGPT-5.2 · high · task 11 · trial 0 · passGPT-5.2 · high · task 24 · trial 1 · passGPT-5.2 · high · task 4 · trial 0 · passGPT-5.2 · high · task 32 · trial 0 · failGPT-5.2 · high · task 20 · trial 0 · passGPT-5.2 · high · task 1 · trial 2 · passGPT-5.2 · high · task 19 · trial 1 · passGPT-5.2 · high · task 0 · trial 2 · passGPT-5.2 · high · task 39 · trial 0 · failGPT-5.2 · high · task 34 · trial 1 · passGPT-5.2 · high · task 31 · trial 0 · passGPT-5.2 · high · task 22 · trial 1 · passGPT-5.2 · high · task 38 · trial 1 · passGPT-5.2 · high · task 46 · trial 2 · passGPT-5.2 · high · task 13 · trial 2 · passGPT-5.2 · high · task 18 · trial 1 · passGPT-5.2 · high · task 39 · trial 1 · failGPT-5.2 · high · task 6 · trial 1 · passGPT-5.2 · high · task 5 · trial 2 · passGPT-5.2 · high · task 19 · trial 0 · passGPT-5.2 · high · task 3 · trial 2 · passGPT-5.2 · high · task 7 · trial 0 · failGPT-5.2 · high · task 47 · trial 1 · passGPT-5.2 · high · task 42 · trial 0 · failGPT-5.2 · high · task 16 · trial 1 · passGPT-5.2 · high · task 15 · trial 1 · passGPT-5.2 · high · task 35 · trial 0 · passGPT-5.2 · high · task 33 · trial 1 · passGPT-5.2 · high · task 47 · trial 0 · passGPT-5.2 · high · task 45 · trial 1 · failGPT-5.2 · high · task 31 · trial 1 · failGPT-5.2 · high · task 9 · trial 2 · passGPT-5.2 · high · task 14 · trial 0 · passGPT-5.2 · high · task 27 · trial 2 · passGPT-5.2 · high · task 4 · trial 2 · passGPT-5.2 · high · task 2 · trial 0 · passGPT-5.2 · high · task 14 · trial 1 · passGPT-5.2 · high · task 38 · trial 0 · passGPT-5.2 · high · task 33 · trial 0 · passGPT-5.2 · high · task 44 · trial 0 · failGPT-5.2 · high · task 0 · trial 3 · passGPT-5.2 · high · task 36 · trial 2 · passGPT-5.2 · high · task 7 · trial 1 · failGPT-5.2 · high · task 44 · trial 1 · failGPT-5.2 · high · task 8 · trial 2 · passGPT-5.2 · high · task 28 · trial 2 · passGPT-5.2 · high · task 11 · trial 2 · passGPT-5.2 · high · task 40 · trial 2 · passGPT-5.2 · high · task 35 · trial 1 · passGPT-5.2 · high · task 24 · trial 2 · failGPT-5.2 · high · task 34 · trial 2 · passGPT-5.2 · high · task 37 · trial 2 · failGPT-5.2 · high · task 49 · trial 2 · passGPT-5.2 · high · task 26 · trial 2 · passGPT-5.2 · high · task 19 · trial 2 · failGPT-5.2 · high · task 25 · trial 2 · passGPT-5.2 · high · task 43 · trial 2 · passGPT-5.2 · high · task 13 · trial 3 · passGPT-5.2 · high · task 1 · trial 3 · passGPT-5.2 · high · task 22 · trial 2 · passGPT-5.2 · high · task 10 · trial 0 · passGPT-5.2 · high · task 30 · trial 2 · passGPT-5.2 · high · task 46 · trial 3 · passGPT-5.2 · high · task 18 · trial 2 · passGPT-5.2 · high · task 21 · trial 1 · passGPT-5.2 · high · task 12 · trial 2 · passGPT-5.2 · high · task 10 · trial 1 · passGPT-5.2 · high · task 41 · trial 2 · passGPT-5.2 · high · task 36 · trial 3 · passGPT-5.2 · high · task 3 · trial 3 · passGPT-5.2 · high · task 9 · trial 3 · passGPT-5.2 · high · task 23 · trial 2 · failGPT-5.2 · high · task 17 · trial 2 · passGPT-5.2 · high · task 32 · trial 1 · passGPT-5.2 · high · task 42 · trial 1 · failGPT-5.2 · high · task 14 · trial 2 · passGPT-5.2 · high · task 48 · trial 2 · passGPT-5.2 · high · task 20 · trial 2 · passGPT-5.2 · high · task 28 · trial 3 · passGPT-5.2 · high · task 31 · trial 2 · passGPT-5.2 · high · task 21 · trial 2 · passGPT-5.2 · high · task 23 · trial 3 · failGPT-5.2 · high · task 39 · trial 2 · failGPT-5.2 · high · task 47 · trial 2 · passGPT-5.2 · high · task 11 · trial 3 · passGPT-5.2 · high · task 7 · trial 2 · failGPT-5.2 · high · task 33 · trial 2 · passGPT-5.2 · high · task 43 · trial 3 · passGPT-5.2 · high · task 38 · trial 2 · passGPT-5.2 · high · task 27 · trial 3 · passGPT-5.2 · high · task 31 · trial 3 · passGPT-5.2 · high · task 49 · trial 3 · passGPT-5.2 · high · task 24 · trial 3 · passGPT-5.2 · high · task 42 · trial 2 · failGPT-5.2 · high · task 29 · trial 2 · failGPT-5.2 · high · task 40 · trial 3 · passGPT-5.2 · high · task 15 · trial 2 · passGPT-5.2 · high · task 6 · trial 2 · passGPT-5.2 · high · task 4 · trial 3 · passGPT-5.2 · high · task 16 · trial 2 · passGPT-5.2 · high · task 26 · trial 3 · passGPT-5.2 · high · task 32 · trial 2 · failGPT-5.2 · high · task 34 · trial 3 · passGPT-5.2 · high · task 20 · trial 3 · passGPT-5.2 · high · task 12 · trial 3 · passGPT-5.2 · high · task 21 · trial 3 · passGPT-5.2 · high · task 25 · trial 3 · passGPT-5.2 · high · task 17 · trial 3 · passGPT-5.2 · high · task 15 · trial 3 · passGPT-5.2 · high · task 5 · trial 3 · passGPT-5.2 · high · task 19 · trial 3 · passGPT-5.2 · high · task 44 · trial 2 · failGPT-5.2 · high · task 30 · trial 3 · passGPT-5.2 · high · task 29 · trial 3 · failGPT-5.2 · high · task 47 · trial 3 · passGPT-5.2 · high · task 8 · trial 3 · passGPT-5.2 · high · task 14 · trial 3 · passGPT-5.2 · high · task 42 · trial 3 · passGPT-5.2 · high · task 35 · trial 2 · passGPT-5.2 · high · task 41 · trial 3 · passGPT-5.2 · high · task 37 · trial 3 · passGPT-5.2 · high · task 22 · trial 3 · passGPT-5.2 · high · task 38 · trial 3 · passGPT-5.2 · high · task 10 · trial 2 · passGPT-5.2 · high · task 16 · trial 3 · passGPT-5.2 · high · task 48 · trial 3 · passGPT-5.2 · high · task 44 · trial 3 · failGPT-5.2 · high · task 2 · trial 2 · passGPT-5.2 · high · task 6 · trial 0 · failGPT-5.2 · high · task 39 · trial 3 · failGPT-5.2 · high · task 32 · trial 3 · failGPT-5.2 · high · task 33 · trial 3 · passGPT-5.2 · high · task 7 · trial 3 · failGPT-5.2 · high · task 45 · trial 2 · passGPT-5.2 · high · task 18 · trial 3 · failGPT-5.2 · high · task 2 · trial 3 · passGPT-5.2 · high · task 35 · trial 3 · passGPT-5.2 · high · task 45 · trial 3 · passGPT-5.2 · high · task 6 · trial 3 · passGPT-5.2 · high · task 10 · trial 3 · passClaude Opus 4.5 · high · task 46 · trial 1 · passClaude Opus 4.5 · high · task 46 · trial 0 · passClaude Opus 4.5 · high · task 36 · trial 1 · passClaude Opus 4.5 · high · task 3 · trial 1 · passClaude Opus 4.5 · high · task 40 · trial 1 · passClaude Opus 4.5 · high · task 36 · trial 0 · passClaude Opus 4.5 · high · task 40 · trial 0 · passClaude Opus 4.5 · high · task 16 · trial 1 · passClaude Opus 4.5 · high · task 12 · trial 0 · passClaude Opus 4.5 · high · task 3 · trial 0 · passClaude Opus 4.5 · high · task 37 · trial 1 · passClaude Opus 4.5 · high · task 29 · trial 0 · failClaude Opus 4.5 · high · task 25 · trial 1 · passClaude Opus 4.5 · high · task 44 · trial 0 · failClaude Opus 4.5 · high · task 18 · trial 1 · passClaude Opus 4.5 · high · task 48 · trial 1 · passClaude Opus 4.5 · high · task 0 · trial 0 · passClaude Opus 4.5 · high · task 41 · trial 1 · passClaude Opus 4.5 · high · task 26 · trial 0 · passClaude Opus 4.5 · high · task 48 · trial 0 · passClaude Opus 4.5 · high · task 26 · trial 1 · passClaude Opus 4.5 · high · task 0 · trial 1 · passClaude Opus 4.5 · high · task 13 · trial 0 · passClaude Opus 4.5 · high · task 13 · trial 1 · passClaude Opus 4.5 · high · task 28 · trial 1 · passClaude Opus 4.5 · high · task 49 · trial 1 · passClaude Opus 4.5 · high · task 8 · trial 1 · passClaude Opus 4.5 · high · task 19 · trial 0 · passClaude Opus 4.5 · high · task 12 · trial 1 · passClaude Opus 4.5 · high · task 49 · trial 0 · passClaude Opus 4.5 · high · task 15 · trial 0 · passClaude Opus 4.5 · high · task 1 · trial 1 · passClaude Opus 4.5 · high · task 19 · trial 1 · passClaude Opus 4.5 · high · task 29 · trial 1 · failClaude Opus 4.5 · high · task 4 · trial 0 · passClaude Opus 4.5 · high · task 1 · trial 0 · passClaude Opus 4.5 · high · task 11 · trial 0 · passClaude Opus 4.5 · high · task 34 · trial 1 · passClaude Opus 4.5 · high · task 17 · trial 0 · passClaude Opus 4.5 · high · task 39 · trial 1 · failClaude Opus 4.5 · high · task 42 · trial 1 · passClaude Opus 4.5 · high · task 28 · trial 0 · passClaude Opus 4.5 · high · task 8 · trial 0 · failClaude Opus 4.5 · high · task 22 · trial 0 · passClaude Opus 4.5 · high · task 22 · trial 1 · passClaude Opus 4.5 · high · task 18 · trial 0 · failClaude Opus 4.5 · high · task 4 · trial 1 · passClaude Opus 4.5 · high · task 43 · trial 0 · passClaude Opus 4.5 · high · task 38 · trial 1 · passClaude Opus 4.5 · high · task 30 · trial 0 · passClaude Opus 4.5 · high · task 30 · trial 1 · passClaude Opus 4.5 · high · task 39 · trial 0 · failClaude Opus 4.5 · high · task 24 · trial 0 · passClaude Opus 4.5 · high · task 38 · trial 0 · passClaude Opus 4.5 · high · task 42 · trial 0 · passClaude Opus 4.5 · high · task 27 · trial 1 · passClaude Opus 4.5 · high · task 9 · trial 1 · passClaude Opus 4.5 · high · task 17 · trial 1 · passClaude Opus 4.5 · high · task 24 · trial 1 · passClaude Opus 4.5 · high · task 43 · trial 1 · passClaude Opus 4.5 · high · task 32 · trial 0 · failClaude Opus 4.5 · high · task 37 · trial 0 · passClaude Opus 4.5 · high · task 16 · trial 0 · passClaude Opus 4.5 · high · task 15 · trial 1 · passClaude Opus 4.5 · high · task 41 · trial 0 · passClaude Opus 4.5 · high · task 11 · trial 1 · passClaude Opus 4.5 · high · task 33 · trial 0 · passClaude Opus 4.5 · high · task 0 · trial 2 · passClaude Opus 4.5 · high · task 25 · trial 0 · passClaude Opus 4.5 · high · task 44 · trial 1 · failClaude Opus 4.5 · high · task 32 · trial 1 · failClaude Opus 4.5 · high · task 20 · trial 1 · passClaude Opus 4.5 · high · task 34 · trial 0 · passClaude Opus 4.5 · high · task 33 · trial 1 · passClaude Opus 4.5 · high · task 14 · trial 0 · passClaude Opus 4.5 · high · task 9 · trial 0 · passClaude Opus 4.5 · high · task 13 · trial 2 · passClaude Opus 4.5 · high · task 20 · trial 0 · passClaude Opus 4.5 · high · task 1 · trial 2 · passClaude Opus 4.5 · high · task 3 · trial 2 · passClaude Opus 4.5 · high · task 5 · trial 1 · passClaude Opus 4.5 · high · task 2 · trial 1 · passClaude Opus 4.5 · high · task 6 · trial 0 · passClaude Opus 4.5 · high · task 45 · trial 0 · passClaude Opus 4.5 · high · task 11 · trial 2 · passClaude Opus 4.5 · high · task 21 · trial 0 · failClaude Opus 4.5 · high · task 4 · trial 2 · passClaude Opus 4.5 · high · task 5 · trial 0 · passClaude Opus 4.5 · high · task 31 · trial 1 · failClaude Opus 4.5 · high · task 28 · trial 2 · passClaude Opus 4.5 · high · task 7 · trial 0 · failClaude Opus 4.5 · high · task 36 · trial 2 · passClaude Opus 4.5 · high · task 35 · trial 1 · passClaude Opus 4.5 · high · task 9 · trial 2 · passClaude Opus 4.5 · high · task 47 · trial 0 · passClaude Opus 4.5 · high · task 31 · trial 0 · failClaude Opus 4.5 · high · task 47 · trial 1 · passClaude Opus 4.5 · high · task 16 · trial 2 · passClaude Opus 4.5 · high · task 46 · trial 2 · passClaude Opus 4.5 · high · task 26 · trial 2 · passClaude Opus 4.5 · high · task 21 · trial 1 · passClaude Opus 4.5 · high · task 17 · trial 2 · passClaude Opus 4.5 · high · task 2 · trial 0 · passClaude Opus 4.5 · high · task 20 · trial 2 · passClaude Opus 4.5 · high · task 40 · trial 2 · passClaude Opus 4.5 · high · task 10 · trial 1 · passClaude Opus 4.5 · high · task 19 · trial 2 · passClaude Opus 4.5 · high · task 39 · trial 2 · failClaude Opus 4.5 · high · task 8 · trial 2 · passClaude Opus 4.5 · high · task 7 · trial 1 · passClaude Opus 4.5 · high · task 15 · trial 2 · passClaude Opus 4.5 · high · task 23 · trial 1 · failClaude Opus 4.5 · high · task 13 · trial 3 · passClaude Opus 4.5 · high · task 7 · trial 2 · failClaude Opus 4.5 · high · task 45 · trial 1 · passClaude Opus 4.5 · high · task 0 · trial 3 · passClaude Opus 4.5 · high · task 14 · trial 1 · passClaude Opus 4.5 · high · task 43 · trial 2 · passClaude Opus 4.5 · high · task 24 · trial 2 · passClaude Opus 4.5 · high · task 3 · trial 3 · passClaude Opus 4.5 · high · task 48 · trial 2 · passClaude Opus 4.5 · high · task 25 · trial 2 · passClaude Opus 4.5 · high · task 49 · trial 2 · passClaude Opus 4.5 · high · task 42 · trial 2 · passClaude Opus 4.5 · high · task 4 · trial 3 · passClaude Opus 4.5 · high · task 30 · trial 2 · passClaude Opus 4.5 · high · task 11 · trial 3 · passClaude Opus 4.5 · high · task 41 · trial 2 · passClaude Opus 4.5 · high · task 12 · trial 3 · passClaude Opus 4.5 · high · task 16 · trial 3 · passClaude Opus 4.5 · high · task 21 · trial 2 · passClaude Opus 4.5 · high · task 14 · trial 2 · failClaude Opus 4.5 · high · task 12 · trial 2 · passClaude Opus 4.5 · high · task 22 · trial 2 · passClaude Opus 4.5 · high · task 23 · trial 0 · failClaude Opus 4.5 · high · task 1 · trial 3 · passClaude Opus 4.5 · high · task 27 · trial 3 · passClaude Opus 4.5 · high · task 44 · trial 2 · failClaude Opus 4.5 · high · task 37 · trial 2 · passClaude Opus 4.5 · high · task 28 · trial 3 · passClaude Opus 4.5 · high · task 10 · trial 2 · passClaude Opus 4.5 · high · task 2 · trial 2 · passClaude Opus 4.5 · high · task 34 · trial 2 · passClaude Opus 4.5 · high · task 36 · trial 3 · passClaude Opus 4.5 · high · task 29 · trial 2 · passClaude Opus 4.5 · high · task 18 · trial 2 · passClaude Opus 4.5 · high · task 8 · trial 3 · passClaude Opus 4.5 · high · task 19 · trial 3 · passClaude Opus 4.5 · high · task 25 · trial 3 · passClaude Opus 4.5 · high · task 46 · trial 3 · passClaude Opus 4.5 · high · task 5 · trial 2 · passClaude Opus 4.5 · high · task 26 · trial 3 · passClaude Opus 4.5 · high · task 22 · trial 3 · passClaude Opus 4.5 · high · task 29 · trial 3 · failClaude Opus 4.5 · high · task 6 · trial 2 · passClaude Opus 4.5 · high · task 40 · trial 3 · passClaude Opus 4.5 · high · task 18 · trial 3 · passClaude Opus 4.5 · high · task 35 · trial 0 · passClaude Opus 4.5 · high · task 9 · trial 3 · passClaude Opus 4.5 · high · task 7 · trial 3 · failClaude Opus 4.5 · high · task 39 · trial 3 · failClaude Opus 4.5 · high · task 30 · trial 3 · passClaude Opus 4.5 · high · task 24 · trial 3 · passClaude Opus 4.5 · high · task 41 · trial 3 · passClaude Opus 4.5 · high · task 21 · trial 3 · passClaude Opus 4.5 · high · task 49 · trial 3 · passClaude Opus 4.5 · high · task 5 · trial 3 · passClaude Opus 4.5 · high · task 15 · trial 3 · failClaude Opus 4.5 · high · task 10 · trial 0 · passClaude Opus 4.5 · high · task 42 · trial 3 · passClaude Opus 4.5 · high · task 48 · trial 3 · passClaude Opus 4.5 · high · task 20 · trial 3 · passClaude Opus 4.5 · high · task 31 · trial 2 · passClaude Opus 4.5 · high · task 14 · trial 3 · passClaude Opus 4.5 · high · task 38 · trial 2 · passClaude Opus 4.5 · high · task 33 · trial 2 · passClaude Opus 4.5 · high · task 32 · trial 2 · passClaude Opus 4.5 · high · task 34 · trial 3 · passClaude Opus 4.5 · high · task 6 · trial 3 · passClaude Opus 4.5 · high · task 47 · trial 2 · passClaude Opus 4.5 · high · task 31 · trial 3 · passClaude Opus 4.5 · high · task 38 · trial 3 · passClaude Opus 4.5 · high · task 45 · trial 2 · passClaude Opus 4.5 · high · task 2 · trial 3 · passClaude Opus 4.5 · high · task 32 · trial 3 · failClaude Opus 4.5 · high · task 43 · trial 3 · passClaude Opus 4.5 · high · task 17 · trial 3 · passClaude Opus 4.5 · high · task 27 · trial 0 · passClaude Opus 4.5 · high · task 27 · trial 2 · passClaude Opus 4.5 · high · task 37 · trial 3 · passClaude Opus 4.5 · high · task 45 · trial 3 · passClaude Opus 4.5 · high · task 44 · trial 3 · failClaude Opus 4.5 · high · task 35 · trial 2 · failClaude Opus 4.5 · high · task 33 · trial 3 · failClaude Opus 4.5 · high · task 47 · trial 3 · passClaude Opus 4.5 · high · task 10 · trial 3 · passClaude Opus 4.5 · high · task 6 · trial 1 · failClaude Opus 4.5 · high · task 35 · trial 3 · failClaude Opus 4.5 · high · task 23 · trial 2 · failClaude Opus 4.5 · high · task 23 · trial 3 · fail
Select any point to inspect its run. Darker points pass; lighter points fail. Points can overlap; the table exposes every run.
Cost / latency coverage: 400/400 · 400/400Cost is upstream model-priced agent spend, not audited billing. Duration is full simulation wall time, including simulated-user and tool work.
Run / modelOutcomeMessagesToolsRepeats
pass310
pass310
pass560
pass690
pass520
pass520
pass680
pass640
pass630
pass640
pass680
pass680
1–12 of 400

Trace inspection · task 46 · trial 1

The conversation. The calls. The results.

Published simulated benchmark data. Read-only historical evidence; nothing here executes tools.
Source provenance, reproduction and what these results establish

This dashboard analyzes complete pinned public benchmark files, with no selection based on outcome. It preserves the publisher’s rewards; Gradia did not execute these models or independently regrade these tasks. Same task specs, trial keys, recorded seeds and shared recorded configuration. Historical runtime dependencies were not independently reproduced. Matching recorded configuration or seeds does not guarantee identical stochastic interactions. Repeated trials on the same task are not independent evidence about all enterprise work.

The rates, gaps, tool coverage and sequences are deterministic calculations. They describe this sample; they are not AI-generated causal explanations, confidence intervals or current model recommendations. Assistant messages measure sequence length, not elapsed time. Repeated calls may be legitimate.

Source: Sierra Research · τ2-bench · recorded cost and performance, MIT license. Upstream commit 2174a603f6d014ef94473ffa95957f6ce27100db. Full source pins and run fingerprints are included in the downloadable JSON report. See the repository generator for reproducible extraction.

From benchmark numbers to your decision

Now ask what “pass” means for your business.

Use your cases and requirements to compare agents, inspect the evidence and identify the next tests. The replay lab below shows why that extra layer matters.

Evaluate your workflow →

Replay lab · Authored evidence tests

Stress-test the evidence.

Keep a public record’s final answer unchanged. Alter its evidence, then inspect exactly which requirement catches the difference.

Public historical traces + deliberately authored replay faults. This replay lab demonstrates Gradia’s checks, not fresh model performance, rankings or independent task success.

01 · Choose a workflow

Delegated agent work

Same answer. Wrong branch.

One authored variant moves a successful tool result outside its required delegated branch. The other marks a delegated run as failed. Both retain the original final answer.

What does the check see?

Switch from the copied answer to the evidence behind it.

PassFailUnknownCounts of replay attempts
Published source replay2 pass / 2 planned
2 pass of 2 planned
Authored comparison variant2 fail / 2 planned
2 fail of 2 planned

A required failure blocks a pass. Missing evidence remains unknown. Select a segment to inspect its meaning.

Exact chart counts and denominators
TRAIL: output preservation and required evidence
ReplayCopied output passesEvidence passEvidence failEvidence unknown
Published source replay2 / 22 / 20 / 20 / 2
Authored comparison variant2 / 20 / 22 / 20 / 2

See the actual structural change

Right result. Wrong place.

Simplified view of one selected replay pair.

Manager
Manager step 1

Required branch

Required search agent
Search agent step 3
Successful web_search resultArguments and result unchanged
Manager step 2

A different step of the same manager

Frozen requirement

The result must belong to the required search-agent branch.

Pass · Result is inside the branch

The successful tool span is a child of the search agent’s step 3.

Final answer: unchangedCopied-output preservation still passes.

This is recorded span containment, not a live execution or a restored application. Other spans and wrapper nodes are omitted. The other TRAIL pair in the chart uses a different authored fault.

Source pair and exact parent change
Selected tool span
f88acd6dcdc7adef
Original parent → authored parent
13ab6b8e76ba8e3e01180c4fc9a8095c
Required agent span
a2e382417dcc3c2b
Frozen requirement fingerprint
adb563df8c64142e585b861233379b94fbe9d3bea6ba69a4121bc8601d28b836
Public source SHA-256
cee9269e6af749801976f88d31ecbcd831b498987fdb98aae07def1af3e5f20c
Verified local replay SHA-256
5042d1a05423d3be9891dab5f421d6f9c8c595768a7fa563e2c3a4ca950a3ff4

02 · Find the requirement

Follow the difference.

Each row is one exact frozen rule. Select it to compare the two sides.

2 of 4 evidence requirements shown
RequirementSourceAuthored variant
1 pass/ 1 applicable1 fail/ 1 applicable
1 pass/ 1 applicable1 fail/ 1 applicable

Several criteria can apply to one attempt. Do not add criterion counts to estimate the number of failed attempts.

03 · Make the next test useful

What would matter in your workflow?

A result from another branch cannot satisfy the work you assigned to this one. Check who did the work and whether the required delegation completed.

A practical next step: Define which delegated branch must produce the result, then preserve that requirement as a regression test.

These are recorded execution trees, not restored application worlds. The changed branch and failure are deliberate demo edits, not newly observed model mistakes.

04 · Understand the limits

A reproducible example.
A bounded claim.

9 selected public records, 18 replay attempts. Across this replay catalog, 18 copied-output checks pass; required evidence records 4 pass, 11 fail and 3 unknown. These are inventory counts across different requirements, not a pooled performance score.

Historical cost and latency are unknown for all 18 replay attempts. This lab draws no cost chart. Its offline qualification made no model or network calls.

What “output pass” means
Typed JSON equality with the copied source output. Demonstrates preservation, not task correctness.
What “evidence pass” means
Outcome under the demo's frozen required evidence assertions. Fail if a required assertion fails; otherwise unknown if required evidence is insufficient; otherwise pass.
What changed
7 of 9 same-answer pairs have different required-evidence patterns. The variants are authored edits, not new observations of the upstream agents.
Reproduction and receipt identity

All replay-lab counts are a curated projection of the offline four-source qualification receipt. Downloading exports this same aggregate data, source pins, definitions, and the selected public span relationships. It contains no raw prompts, tool payloads, screenshots or customer records.

docs/examples/evals-public-suite/insights-qualification.json

Source receipt SHA-256

db5c5b9531dbe09c798adc750531fd340d6aa307d04fef0e9c53194a15dc8ca4

Replay-lab sources

Attribution is not endorsement. Original materials retain their publishers’ MIT license notices. This replay lab redistributes aggregate measurements, selected span relationships and authored explanations only.

  • TRAIL Patronus AI · MIT
    Source pin0ffbed9db859b4a66250dc783fa4dccf86869595
  • τ-bench Sierra · MIT
    Source pin59a200c6d575d595120f1cb70fea53cef0632f6b
  • OpenCUA The OpenCUA Authors · MIT
    Source pindfc91ba89f700d10f26ec50362d308571482ab8b
  • AgentNet The OpenCUA Authors · MIT
    Source pind76ee50a63fad81cfdbe576416757d7c2091ed50

Now bring your work

Which agent would you
trust with this workflow?

Bring your agent or traces, one workflow, and the decision you need to make. Agree the requirements, evaluate the evidence, and leave with a reviewable report and the next tests to run.

Plan an evaluation pilot