What Lilian Weng thinks about LLMs
Machine-learning researcher; has written the Lil'Log survey posts on how a model technique works since 2017, and worked at OpenAI from 2018 to 2024.
Everything they publish, on ppll ↗
Lilian Weng did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
4 dated positions, 2024 to 2025, in their own words. Our reading of what Lilian Weng has said — not written or endorsed by them.
-
Our readingLarge language models cannot reliably self-correct their own mistakes without external feedback.
Their wordsthis self-correction capability turns out to not exist intrinsically among LLMs and does not easily work out of the box, due to various failure modes
-
Their wordsmodel CoTs could be biased due to lack of explicit training objectives aimed at encouraging faithful reasoning
- 10 months earlier
-
Their wordsTo avoid hallucination, LLMs need to be (1) factual and (2) acknowledge not knowing the answer when applicable.
↗Extrinsic Hallucinations in LLMslilianweng.github.io 1st of 2 in this piece
-
Their wordsThese empirical results from Gekhman et al. (2024) point out the risk of using supervised fine-tuning for updating LLMs' knowledge.
↗Extrinsic Hallucinations in LLMslilianweng.github.io 2nd of 2 in this piece