What Lilian Weng thinks about design
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024.
Everything they publish, on ppll ↗
Lilian Weng did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
3 dated positions, 2024, in their own words. Our reading of what Lilian Weng has said — not written or endorsed by them.
-
Their wordsReward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
↗Reward Hacking in Reinforcement Learninglilianweng.github.io 1st of 2 in this piece
-
Our readingMore capable AI agents are more likely to find and exploit flaws in their reward functions.
Their wordsA more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards.
↗Reward Hacking in Reinforcement Learninglilianweng.github.io 2nd of 2 in this piece
- 8 months earlier
-
Their wordsIt has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model.
↗Diffusion Models for Video Generationlilianweng.github.io 1st of 2 in this piece