korrents

korrents · piece

Reward Hacking in Reinforcement Learning

Lilian Weng · 28 Nov 2024 · lilianweng.github.io

2 korrents from this piece

Lilian Weng did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
  2. A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards.