korrents

← Mental models

When a measure becomes a target

Reward a number and people, or machines, will find ways to move the number without the thing it measured (Goodhart's law).

Our reading of the idea · 8 claims from 6 people, in their own words

Who thinks with this

Ajeya CotraLilian WengHenrik KarlssonRyan GreenblattScott AlexanderZvi Mowshowitz

It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning.

  1. Lilian Weng Machine-learning researcher Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function. Reward Hacking in Reinforcement Learninglilianweng.github.io · 28 Nov 2024All korrents from this piece
    Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.

The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.

  1. Ajeya Cotra AI researcher and grantmaker Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place. Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com · 1 Sept 2026 · 1:56:09 into the videoAll korrents from this video
    Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.

    Watch from 1:56:09 plays here

    ↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com

    1 Sept 2026 · video · 2h 20m · spoken · machine transcript

Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.

  1. Zvi Mowshowitz Writes Don't Worry About the Vase At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments. HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com · 31 Aug 2026
    At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.

If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.

  1. Ryan Greenblatt Chief scientist at Redwood Research One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out Ryan Greenblatt – What happens once AI can automate AI research?youtube.com · 11 Aug 2026 · 1:40:37 into the videoAll korrents from this video
    One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out

    Watch from 1:40:37 plays here

    ↗Ryan Greenblatt – What happens once AI can automate AI research?youtube.com

    11 Aug 2026 · video · 2h 12m · spoken · machine transcript

A feedback loop that reinforces a behavior pulls it toward whatever improves the loop's own score.

  1. Henrik Karlsson Swedish essayist who writes Escaping Flatland If there is a feedback loop reinforcing a behavior, it is hard not to go in the direction that improves the score. First we shape our feedback loops; then they shape ushenrikkarlsson.xyz · 26 Aug 2026All korrents from this piece
    If there is a feedback loop reinforcing a behavior, it is hard not to go in the direction that improves the score.

Reward hacking has stopped being myopic: these agents were willing to embark on cheating projects that would take weeks to pay off.

  1. Ajeya Cotra AI researcher and grantmaker So it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat. Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com · 1 Sept 2026 · 1:02:59 into the videoAll korrents from this video
    So it seemed like they were willing to embark on quests that might take weeks to succeed um in order to cheat.

    Watch from 1:02:59 plays here

    ↗Ajeya Cotra – "This might be the clearest warning shot we ever get"youtube.com

    1 Sept 2026 · video · 2h 20m · spoken · machine transcript

More capable AI agents are more likely to find and exploit flaws in their reward functions.

  1. Lilian Weng Machine-learning researcher A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards. Reward Hacking in Reinforcement Learninglilianweng.github.io · 28 Nov 2024All korrents from this piece
    A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards.

AI is positively reinforced for success on benchmarks, including impossible ones, and then negatively reinforced for getting caught cheating.

  1. Scott Alexander Essayist; writes Astral Codex Ten We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating. Nicholas Decker In Hellastralcodexten.com · 1 Sept 2026
    We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.

More in systems