AI agent benchmarks

A benchmark is a frozen measurement made inside a governed Gradia Universe: exact tasks, evaluator, model panel, results, and release evidence. One Universe can support multiple benchmark editions; every public edition should point back to the world it measured.

Explore the worlds behind the measurements →

Have a benchmark you need to trust?

Benchmark Audit examines the quality of the measurement: task clarity, grader behavior, evidence gaps and findings that need human review. Start with the benchmark you already use.

AI benchmark auditing: method, deliverables and limits →

Certified editions

Signed release required

No certified benchmark has been released

Research instruments appear below. This catalogue stays empty until an exact benchmark edition carries a frozen environment fingerprint, admitted evaluator, eligible results, and signed certificate.

Research benchmarks

Public instruments, before certification

These studies are citable and inspectable. They do not inherit a Gradia certificate merely by appearing here. Training experiments, methods papers, and product explainers stay on Research unless they define an evaluation instrument.

Published metaevaluation study · DOI-backed

Environment: Branchable evaluator-stress environment

Reward-Hacking Wind Tunnel

An oracle-witnessed stress test of benchmark scorers, causal localization, evaluator repair, and held-out re-attack.

Boundary: This measures scorer gameability and repair behavior, not model capability ranking or measured post-training improvement. It is public research, not a certified Gradia benchmark edition.

Preliminary machine results · human adjudication pending

Environment: Branchable synthetic underwriting environment

Conditionally Approved

Long-horizon evaluation under changing evidence, authority, and time, with gradable attempts separated from infrastructure and provider exclusions.

Boundary: A citable research benchmark, not a certified public Gradia benchmark edition. Provider ranking and model-attributed failure claims remain withheld.

Instrument released · agreement study pending

Environment: Standalone synthetic sales environment

The Value Engine Benchmark

A methodology-controlled synthetic enterprise-sales negotiation benchmark with deterministic evidence grading and a frozen multi-model study design.

Boundary: Public research outside the certificate-backed Gradia catalogue. Gradia ingestion, evaluator admission, and public certification remain separate gates.

AI agent benchmarks and evaluation research — Gradia