It is fundamentally difficult to design a reward function that accurately captures the intended goal in reinforcement learning.
Lilian Weng Machine-learning researcher Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function. Reward Hacking in Reinforcement Learninglilianweng.github.io · 28 Nov 2024All korrents from this piece
Their wordsReward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.