Benchmark audit

AI benchmark quality audit

Before relying on a benchmark score, examine the measurement that produced it. Gradia Benchmark Audit helps your team inspect an exact benchmark edition, review its evidence and decide which claims can be supported.

Have a workspace? Open Console. New managed workspaces require an invitation.

Who it helps
AI evaluation leads, model governance teams, benchmark owners and enterprise buyers comparing agent claims. Start when a score is influencing a procurement, release or research decision.
Availability and scope
The Console supports evidence registration, isolated inventory, findings, independent review and benchmark-quality reporting. External harness execution and calibration require an approved isolated runner and exact machine receipts; uploading a repository does not make it executable.

Work with Gradia

Scope a benchmark audit

A focused, assisted engagement. No Console account needed to start the conversation.

Review one benchmark edition and the decision it needs to support. Agree which checks can run, how findings will be reviewed and what evidence you will receive.

Bring
One workflow, the decision you face, a person who knows the work and a description of the evidence available.
Receive
A scoped proposal with deliverables, required inputs, supported checks, review responsibilities, timing and fees. Work begins after those are agreed.
Email Gradia to scope a pilot →

Opens your email app. You can also write to rudy@gradiahq.com. Describe the workflow first; keep confidential records for the agreed transfer process.

How do we know whether an AI benchmark score is trustworthy?

Check the benchmark, its harness and its evidence against the decision you want to make. Gradia preserves separate findings across fourteen quality dimensions, including task clarity, grader validity, realism, provenance, reproducibility and statistical fitness. There is no overall trust score: a passed dimension cannot erase a failed or unassessable one.

How the engagement works

  1. State the claim and freeze the source

    Identify the exact benchmark edition or repository commit, source digest, rights basis and intended use. State the claim narrowly enough to test. Public availability alone does not establish permission for training, derivative work or publication.

  2. Inventory the benchmark and its gaps

    Gradia can inspect an approved repository, immutable object, project upload or existing Gradia environment. The isolated inventory examines data without running the submitted code. It records candidate tasks and graders as discovery hints, with explicit reconstruction gaps that must be resolved before execution claims are supported.

  3. Test the measurement with the right evidence

    Freeze an audit policy and evaluation scope. Use static probes and, for supported frozen Gradia environments, native probes. External harness and calibration work need approved runner receipts binding the planned work, completed cells and artifacts. Missing work stays visible rather than becoming a passing result.

  4. Review findings and decide what to change

    A different authorized person reviews each finding and its interpretation limits. A complete report requires the applicable evaluation, calibration and remediation reviews. Preserve unresolved issues, issue a report with its eligibility status and compare later editions without rewriting the earlier evidence.

What you leave with

A benchmark-quality profile
Read each dimension's verdict, evidence and limitation separately. Unassessable findings explain why the available evidence cannot support a conclusion.
A reviewed reporting artifact
When prerequisites are complete, issue a signed internal report and a shareable projection that omits unsafe evidence references. Free-form prose still needs confidentiality review before sharing.
A traceable remediation history
Record decisions to accept, remediate, regenerate or retire findings. Later signed re-audit comparisons show changed evidence and verdicts; they do not invent the reason a result changed.

Illustrative example

A successful command is not a completed evaluation

Imagine a synthetic runner that exits successfully but skips part of its planned task list. The exit code alone looks reassuring. An audit compares the expected work with the receipt and required artifacts, keeping the incomplete result ineligible for a complete-run claim until the discrepancy is resolved.

Synthetic illustration of an execution-evidence check. This is not a finding about a customer, model or public benchmark.

Before you start

Does a Benchmark Audit certify that our AI model is compliant?
No. It examines benchmark quality and the claims the evidence supports. It is not a regulatory compliance certificate or an ordinary agent-performance grade. A report signature establishes authorship and byte integrity; it does not turn an ineligible result into a pass.
Can we bring an existing open-source benchmark?
Yes, subject to an exact source version, supported intake and a rights basis covering the intended use. Inventory can expose missing runtime or evaluation contracts. Executing an external harness remains a separate reconstruction and runner qualification step.
Does a low pass rate prove a task is hard?
A low pass rate can also reflect a broken harness or invalid grader. Gradia requires evaluation validity before interpreting calibration, then preserves the exact model and scaffold panel, sample counts and uncertainty. Difficulty claims remain scoped to that declared panel.

Keep the next step connected

Request a benchmark audit

Bring the exact benchmark edition and the decision its score should support.

Inspect Gradia's research

Explore the research instruments, methods and stated limitations.

Explore the benchmark catalogue

See how published editions and research instruments are distinguished.