korrents

← People and their mental models

Lilian Weng's mental models

2 claims Lilian Weng made fit 1 mental model. Most often: When a measure becomes a target. Everything they said here.

Models they name

Their own words name the idea.

When a measure becomes a target

Reward a number and people, or machines, will find ways to move the number without the thing it measured (Goodhart's law).

It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning.

  1. Lilian Weng Machine-learning researcher Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function. Reward Hacking in Reinforcement Learninglilianweng.github.io · 28 Nov 2024All korrents from this piece
    Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.

More capable AI agents are more likely to find and exploit flaws in their reward functions.

  1. Lilian Weng Machine-learning researcher A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards. Reward Hacking in Reinforcement Learninglilianweng.github.io · 28 Nov 2024All korrents from this piece
    A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards.