09 · The living Scorer Gameability Index
One recomputable number per benchmark scorer
The instrument distills into standing infrastructure: a per-benchmark gameability index — oracle-witnessed exploit density under the canonical adversary, with a bootstrap CI, pinned to the dataset snapshot, judge identity, catalog version, and git SHA. Each index card is content-addressed; anyone recomputes the number offline without trusting the authors.
The ordering is mechanistic, not difficulty-driven: gameability tracks the judge’s ability to independently verify, not the benchmark’s difficulty for models. GSM8K and HLE sit at opposite ends of both scales.
Oracle-witnessed exploit density, real claude-opus-5 judge
L1 metamorphic (light) · L2 adaptive search (solid) · 95% item-clustered bootstrap CIs
Adaptive budget roughly doubles density on every leaking scorer and cracks the near-resistant key-match scorer (GPQA-Diamond 6.6 → 15.2) — but opens no gap where none exists. HumanEval and GSM8K sit on the independent-verification frontier at exactly zero. SWE-bench Verified is a structurally different archetype — a weak test grader attacked by test-gaming patches, measured live in Docker — so it ships as its own card and is never pooled with the LLM-judge rows.
The agentic archetype: reward-hacking the environment checker
5 typed reward-hacks × 5 weak-checker archetypes — a clean diagonal (exploits /1k)
When the scorer grades a final environment state (Terminal-Bench, OSWorld, SWE-bench), the checker is a proxy — and a proxy is exactly what a reward-hacking policy games. Each of five typed hacks games exactly its target checker and no other (100 oracle-witnessed exploits over 700 attempts; both controls at 0), every one witnessed-localized to the single observable variable it manipulates. Each weak checker has one dominant reward-hack — a benchmark’s blind spot is specific, not uniform.
The archetype in the wild: SWE-bench Verified, live in Docker (N=100)
test-gaming patches vs. three weak graders and the real hidden-suite oracle (exploits /1k)
150test_skip_injection vs present-tests grader
143.3exit_zero_wrapper vs exit-code grader
3.3assertion_defang
0public_subset_gaming in the wild
0strong oracle: restored hidden suite catches every attack
0 / 0controls: gold_fix & no_op both clean
One container run per (instance, patch), read through three proxy lenses: 1,800 attempts, 89 oracle-witnessed exploits, 49.4/1k overall [CI 36.7–62.8]. Every weak proxy that reads the tree “as the patch left it” leaks; the official harness, which restores the dataset’s original tests, leaks nothing. For benchmarks with no oracle at all, a second regime (metamorphic invariance, bias probes, judge-disagreement geometry) carries the audit with honestly weaker verdicts.