Paper 01
2 Sep 2026
Version 1.0.2
Published research artifact · one-seed paired GRPO
Reward Hacking in the RL Loop: Oracle-Witnessed Localization and Repair of Reward-Model Exploits During Training
A training-time study of what happens when an optimizer repeatedly probes a reward channel for exploitable seams. Matched 300-step GRPO/LoRA arms preserve exact manifests, frame chains, final adapters, analysis, and model-backed replay receipts so proxy success and oracle truth remain separately inspectable.
- E1
- Matched 300-step GRPO/LoRA arms · two verified 313-frame chains
- E2
- RLVR control gap 0 · gameable arm 58/64 proxy vs 1/64 oracle · gap 0.890625
- E3
- 55 property/control checks · exact paired analysis · two model-backed final replays
Claim boundary: The frozen result is one seed on Qwen2.5-0.5B-Instruct and GSM8K. Both arms lost oracle accuracy, so the study establishes reward-channel exploitation and amplification in this cell—not capability improvement, population prevalence, a frontier-model result, or production-reward behavior. Repeated-seed variance and real-policy localization/repair remain open.
doi:10.5281/zenodo.22259605