korrents

← People and their mental models

Zvi Mowshowitz's mental models

2 claims Zvi Mowshowitz made fit 2 mental models. Most often: Second-order thinking, When a measure becomes a target. Everything they said here.

Models they name

Their own words name the idea.

When a measure becomes a target

Reward a number and people, or machines, will find ways to move the number without the thing it measured (Goodhart's law).

Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.

  1. Zvi Mowshowitz Writes Don't Worry About the Vase At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments. HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com · 31 Aug 2026
    At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.

Models we see in what they say

Our reading: their claim applies the idea without naming it. The claim is theirs; filing it here is ours.

Second-order thinking

Ask "and then what?": follow a decision past its first effect to the ones that come after.