The Value Engine · Gradia Research
An interactive explainer · after the paper by Rudy M. Celekli

The Value Engine Benchmark

Can an AI agent actually sell — and can you audit the receipts?

LLM agents are being sent into negotiations, sales cycles, and advocacy — conversations where the other side holds private information and only partly shares your goals. Yet nearly every agent benchmark tests the opposite: single-shot tasks in environments that want to be helped. The Value Engine Benchmark (VEB) closes that gap with a six-week simulated enterprise-sales deal, a stochastic buying committee that makes the agent earn every fact, and a judge that must quote the transcript for every point it awards.

0
graded trajectories
0
model endpoints
0
scenarios · 9 industries
0
evidence-cited features / run
01 · Why negotiation breaks benchmarks

The environment stops being your friend

SWE-bench hands you a codebase. WebArena hands you a website. τ-bench hands you a user who wants help. A buying committee hands you nothing.

When an agent negotiates or sells, its counterpart is strategically non-cooperative: it holds private information, has goals only partly aligned with the agent’s, and reveals value only in response to earned trust. Success is not one verifiable end state but a trajectory of earned commitments — a quantified pain, a meeting with the person who controls budget, a price held under pressure. The paper argues existing benchmarks sidestep three difficulties this regime creates.

Difficulty 1

Partial information, gated disclosure

The environment must withhold facts until the agent earns them through the right kind of question. A benchmark that hands over the answer on request can’t measure discovery skill — only discovery effort.

Difficulty 2

Process credit, not just outcome

A deal can close for the wrong reasons (an unforced discount) or fail despite excellent execution. A scalar win/loss reward conflates skill with luck — and rewards value-destroying shortcuts.

Difficulty 3

Methodology attribution

Does a selling methodology actually improve outcomes? Isolating its effect means holding model, scenario, and counterpart fixed while varying only the method — a control almost never built into agentic benchmarks.

Typical agent benchmark

Cooperative or static environment · single verifiable end state · scalar reward · one deterministic interlocutor you can overfit.

VEB

Adversarial stochastic committee · trajectory of earned sub-goals · evidence-cited scorecard · buyer re-sampled every seed.

02 · The environment

The room you have to win

Each of VEB’s nine scenarios is a realistic B2B deal: a committee of LLM-simulated personas — a champion, an economic buyer, technical gatekeepers, procurement — each with a personality, private notes, and gated facts that unlock only when the right conversational move lands. The seller works a bounded six-week calendar against a genuine deadline, an incumbent competitor, and scripted mid-cycle pressure events. Simply asking for a fact does nothing; the simulator releases it only when the seller executes the qualifying move.

Watch one discovery arc play out:

The buying committee · facts are earned, not asked for

SELLER the agent under test CH Champion · opened the door EB locked EB Economic buyer · controls budget GK Technical gatekeeper · can veto PR Procurement · enters mid-cycle INCUMBENT quantifying_question 🔒 gated fact · quantified pain [t·9] “$1.2M a year in missed SLAs.” champion_test [t·41] “If the pilot clears 98%, I’ll sign before renewal.”
Week 1 — only the champion will talk. Every fact of substance sits behind a gate; the economic buyer will not take the meeting.

Gate types include quantifying_question, process_question, champion_test, competitor_probe, and eb_meeting_held. Quotes here are illustrative of the mechanic; scenario facts are defined per deal. Scenarios are admitted only if a weak scripted seller loses and a disciplined reference seller wins — so “impossible deal” can never masquerade as “model failure.”

The calendar imposes genuine opportunity cost — a slot spent on the eager champion is a slot not spent earning the economic-buyer meeting — and the seller may also plan privately between touches, producing internal artifacts (a value model, a stakeholder map) that are graded even if never sent. Difficulty comes in three tiers: warm winnable deals that test disciplined closing, mid-tier gauntlets with veto-holding gatekeepers, and adversarial hard deals — like a post-breach retailer where the motivated CISO doesn’t control budget and urgency is a trap.

03 · Evidence-graded rewards

Quote it, or it didn’t happen

VEB’s first signature choice: the reward is not a scalar, it’s an audit trail. An LLM judge reads the full transcript and fills a structured qualification scorecard — a full MEDDPICC assessment, a Three-Whys root-cause check, economic-buyer engagement, mutual-action-plan completion, champion strength, and price integrity. Every score above zero must cite a verbatim, turn-indexed transcript quote. No citation, and the harness reduces the score to its floor. A human reviewer can check any sub-score against the exact evidence the judge relied on.

One episode’s scorecard, assembling

M
Metrics
E
Economic buyer
D
Decision criteria
D
Decision process
P
Paper process
I
Identified pain
C
Champion
C
Competition

Each letter is scored 0–3 against fixed anchors (0 unknown · 1 assumed · 2 confirmed by one source · 3 corroborated with evidence) — and the lowest letter is the deal’s real score. This illustrative deal has a paper-process hole.

Every filled pip snaps to evidence like  [t·23] “Legal needs the DPA signed before any pilot data moves.”  — the judge’s rule is quote or abstain: a pillar with no supporting transcript evidence is scored at its floor, never inferred.

Deal Velocity Index — weights encode where deals die

MEDDPICC · 40 3 Whys · 20 EB · 15 MAP · 15 CH · 10
≥75 commit-eligible50–74 develop<50 rebuild

Sale Quality Score — process outweighs the close

SQS=0.6 × DVI+20 × PriceIntegrity+Outcome won 20 · no-decision 6 · lost 0
process quality · 60 price integrity · 20 outcome · 20
80 of 100 points reward how the sale was runonly 20 reward whether it closed

A win that gifts a large discount forfeits price-integrity points and typically depresses DVI — so a value-destroying close scores below a disciplined no-decision. “Won/lost” itself is a declarative predicate over the transcript (required facts elicited in the buyer’s words, trust threshold, real EB meeting, confirmed MAP dates, discount within tolerance) — not judge opinion. MEDDPICC scores shown are an illustrative episode.

04 · Methodology-controlled tracks

One variable, everything else nailed down

The second signature choice: every scenario runs as a matched pair. The out-of-box (OOB) seller gets only the deal brief; the pack seller is the identical model, on identical seeds, against the identical stochastic buyer, plus one thing more — an explicit sales methodology as a system prompt. Because only the methodology varies, the paired difference is a near-causal estimate of methodology lift.

Same model, same seed — does the playbook help?

Track A · out-of-box  seed 07

Deal brief + environment interface. Nothing else.

Track B · + methodology pack  seed 07

Same brief, same buyer, same dice — plus the Value Engine methodology in the system prompt.

0 +3 −3 +3.18 Easy 780 pairs · +4.1pp wins −1.44 Mid 585 pairs · +0.3pp wins +0.01 Hard 390 pairs · +2.8pp wins +0.93 Pooled 1755 pairs · +2.6pp wins ΔSQS, pack − OOB · whiskers = 95% CI, seed-matched paired bootstrap (5,000 resamples)

The pack lifts pooled clean-win rate by +2.6pp, concentrated in winnable deals: +4.1pp on easy, only +0.3pp on mid, +2.8pp on hard. The SQS lift is significant on easy (+3.18), significantly negative on mid (−1.44) — where the pack’s prescribed sequence consumes turns the scenario won’t reward — and indistinguishable from zero on hard.

The honest headline: a prompt-injected playbook raises the floor on execution but cannot manufacture access or budget the scenario withholds. Mechanically, the lift appears exactly where behavior moves — economic-buyer attendance rises 28.4% → 31.1% and MAP completion 43.6% → 44.2%. Models that convert the pack into behavior gain; models that convert it only into vocabulary do not.

05 · The released grid

3,510 deals, graded and shipped

The frozen evaluation grid is the full cross-product — every model on every scenario, at uniform seed depth, on both tracks — costing 2.48 billion seller-side tokens and $19,189 to run. Every transcript, scorecard, seed, and cost is released, plus an enrichment corpus of 17 evidence-cited behavioral features per trajectory.

The grid · 9 scenarios × 13 models × 15 seeds × 2 tracks

9 × 13 × 15 × 2 =
0
graded cells · each mosaic tile = one model×scenario pair (30 trajectories) · 98.5% carry a full LLM-judge grade; the 1.5% heuristic fallbacks are flagged in the release

Leaderboard · mean SQS, pooled over both tracks

13 endpoints from 6 labs; green bars mark the two open-weight models (kimi-k3, inkling†), full members of the grid. Beyond the top two OpenAI endpoints the field compresses into a 51–66 band separated more by price discipline than closing power.

gemini-3.5-flash · the win-rate mirage

3rd-highest OOB clean-win rate (11.8%) yet 11th on SQS. Its wins hold price at 0% discount — but across all episodes it concedes a 23.1% mean discount, collapsing price integrity to 0.46. It either closes clean or gives the store away.

claude-fable-5 · the conversion gap

High SQS on process quality (68.5 pack) while almost never producing a clean win (0.7% OOB, 7.4% pack): it qualifies rigorously but fails to convert. A raw win-rate leaderboard would hide both pathologies.

The behavioral decomposition tracks the leaderboard almost exactly: the single strongest correlate of rank is economic-buyer engagement. The leader reaches the EB with a conditional commitment on 71.9% of pack runs and confirms 64.4% of plan dates; the bottom model manages 5.2% and 22.0%. Closing power is downstream of earning access — not a substitute for it.

06 · Cost, quality, speed

The frontier collapses to a single point

Cost isn’t estimated — every model call’s tokens and realized USD spend are logged per cell. Plotting sale quality against cost per completed deal, the trade-off practitioners expect simply isn’t there: gpt-5.6-sol is simultaneously the highest-quality and the cheapest-per-deal seller on the board, delivering 30.5 SQS points per dollar and strictly dominating every other endpoint.

Two ways to count the spread — both log scales

Cost per completed deal $1 $10 $100 $1000 gpt-5.6-sol · $2.38 SQS 72.5 · best AND cheapest $579.66 rests on only 11 wins 244× spread Cost per run (denominator-free check) $0.10 $1 $10 $100 inkling† $0.30 gpt-5.6-sol $1.00 claude-fable-5 $23.62 · ~78×
30.5SQS points per USD for the frontier leader
244×per-deal cost spread — with no quality trade-off
~78×spread survives the denominator-free per-run check

Caveats the paper states itself: per-deal costs for weak closers divide by small win counts (the $579.66 point rests on 11 wins), and gpt-5.6-sol / grok-4.5 pricing assumes the provider’s adjacent-tier list rate pending published prices. The qualitative conclusion is robust to both. Expensive endpoints are not buying quality — most of the field is Pareto-dominated on price and quality together.

07 · Robustness & reliability

Interrogating the judge

An LLM-graded benchmark stands or falls on its judge. VEB’s judge of record (claude-sonnet-4-6, temperature 0) is constrained to quote-or-abstain, and the roll-up math is deterministic — the judge scores atomic fields, never the headline number. But how much of the leaderboard is judge latitude?

The hardened re-grade · remove the judge’s discretion, watch what moves

−4 −3 −2 −1 0 +1 Δ mean SQS per model, hardened judge vs default (2,969 re-graded cells, 11 closed-weight models) claude-fable-5 −3.05 leaned hardest on judge generosity mean −0.52 slight upticks opus-4-6, grok-4.20 Top two of the leaderboard: unchanged.

The hardened judge reconstructs the objective dimensions — MEDDPICC coverage, EB attainment, MAP dates — deterministically from the turn-indexed event log, leaving the LLM only the irreducibly subjective calls (trust, price integrity, champion quality). Mean SQS moves just −0.52; the ordering survives an adversarial re-grade. The re-grade predates the two open-weight arms and covers the 11 closed-weight models, a scope the paper reports explicitly.

Read this before quoting r = 0.894

r = 0.894 is the agreement between the cost-reduced enrichment ensemble and a frontier reference panel on 139 doubly-graded trajectories (mean absolute difference 0.047 on a 0–1 scale; categorical labels match on 89.9–98.6% of cases).

It is a model–model calibration, not a validation against human experts. The paper is explicit: a pre-registered human-expert study — weighted Cohen’s κ and Krippendorff’s α against the frozen grid, with a blind arbitration subset and override audits — is forthcoming, and that is the load-bearing reliability evidence. Until it lands, all published numbers are single-judge, same-family-exposed grades.

The paper is unusually forthcoming about its own weak points:

  • Same-family judge confound. The judge of record is an Anthropic model, and the two largest methodology-pack SQS lifts belong to Anthropic sellers (+5.4 sonnet-4-6, +5.3 fable-5). The pack-lift ranking is flagged as provisional until the cross-family panel re-grade and family-bias audit are published.
  • Simulated buyers are not human buyers. An LLM committee may be exploitable in ways a real executive is not; sim-to-real transfer is unestablished future work.
  • Conflict of interest, stated plainly. VEB is built by the authors of the methodology it evaluates. The remedy offered: everything released, deterministic inspectable math, and results that undercut the sales pitch (a negative mid-tier lift, a small pooled effect).
  • Scenario admission is a selection effect. Scenarios are admitted only if a reference seller running this methodology can win them — so the lift measures “does the methodology help on methodology-winnable deals,” not “does it beat alternatives on arbitrary deals.”
  • Wins concentrate on few behaviors. Reaching the EB and holding price predict most clean wins; the leaderboard partially measures whether a model triggers the behaviors the simulator rewards.
08 · Beyond the leaderboard

A benchmark that doubles as a gym

Because the buyer is re-sampled per seed, VEB resists frozen-replay gaming — and produces something valuable as a by-product: seed-matched win/loss pairs. Hold model, scenario, and track fixed, vary only the buyer seed, and 43.2% of cells contain both a clean win and a non-win: episodes identical in every controllable dimension that diverged only on the buyer’s stochastic choices. That is precisely the substrate DPO, reward modeling, and process reward models consume.

Is there a training signal? The reward is non-degenerate

r = 0reward = SQS / 100r = 1

Illustrative distribution shape — neither an all-fail spike nor a solved-at-1 pile-up. Grand-mean reward 0.60; even the leader clears the SQS ≥ 60 bar on only 67.8% of rollouts, leaving headroom.

0.75
mean normalized entropy of the per-model reward histogram (10 bins) · range 0.63–0.84
43.2%of cells hold a matched win / non-win pair
12.1mean intra-cell SQS standard deviation
0.60grand-mean per-rollout reward

Dense enough to separate policies, hard enough to leave room to improve. Next steps already scaffolded: the human-expert κ/α addendum, a multi-methodology track, and on-environment fine-tuning of the two open-weight models (kimi-k3, inkling) — with the hardened judge standing guard as an adversarial reward model against reward hacking.

The closing aspiration is worth quoting in spirit: the community should move from measuring whether agents produce sales-shaped text to measuring whether they actually sell — earning the meeting, quantifying the pain in the buyer’s own words, and defending value with integrity. And because the scoring penalizes invented evidence, coercive discounting, and value-destroying closes by construction, an agent that manipulates rather than qualifies scores poorly — an evaluation design doing double duty as a safety posture.