What Nathan Lambert thinks about reinforcement learning
Research scientist at the Allen Institute for AI, where he works on post-training open language models, and the author of the AI blog Interconnects.
Nathan Lambert did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
6 dated positions, 2025, in their own words. Our reading of what Nathan Lambert has said — not written or endorsed by them.
-
Their wordsThere's not many worlds where China cannot train AI models. I think export controls are decapping the amount of compute or the density of compute that China can have.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 6th of 44 in this recording
-
Their wordsThere's some research that shows that the distribution is actually the limiting factor. So language models haven't yet made misinformation particularly change the equation there.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 9th of 44 in this recording
-
Their wordsthese open models are probably going to keep coming for the time being, whether or not we want to stop them, and stopping them might make it even worse and harder to prepare.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 19th of 44 in this recording
-
Their wordsAnd the important thing to say is that no matter how you want the model to behave, these RLHF and preference-tuning techniques also improve performance. So, on things like math evals and code evals, there is something innate to these, what is called contrastive loss functions.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 22nd of 44 in this recording
-
Their wordsAnd these reasoning behaviors emerge naturally. So these things like, "Wait, let me see. Wait, let me check this. Oh, that might be a mistake." And they emerge from only having questions and answers.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 23rd of 44 in this recording
-
Their wordsAnd humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.
↗DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459youtube.com 41st of 44 in this recording