Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
Zvi Mowshowitz Writes Don't Worry About the Vase At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments. HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com · 31 Aug 2026
Their wordsAt the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.