scaling laws
Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.
FilterEveryone, all time
- AR
Armin Ronacher quoted
Our readingAI model capability progress is not going to slow down soon
Their wordsSo far we haven't really slowed down with what these models do and I don't think it's going to stop right now.
- 7 days earlier
- FC
François Chollet quoted
Their wordsIn verifiable domains, model capability scaling should remain unbounded. Models will simply keep improving by "absorbing more and more of the computational universe", which is infinite by construction.
- 1 day earlier
- DH
David Heinemeier Hansson quoted
Their wordsSo we should also have the humility. As amazing as the LLMs are now, it could be that they eventually plateau. We haven't seen any evidence of it yet, and I think this is also why we're seeing this absolute gobsmacking levels of investment, because so far the scaling laws are true, and the more billions are poured in, the more intelligence comes out.
↗DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux | Lex Fridman Podcast #501youtube.com 35th of 41 in this recording
- 3 weeks earlier
- FC
François Chollet quoted
Our reading · no longer heldThe LLM line of research will reach a capability plateau.
Recalls holdingIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
SaidIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
- FC
François Chollet quoted
Their wordsI believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
- 5 months earlier
- JH
Jensen Huang quoted
Their wordsThe amount of data that we use to train models is going to continue to scale to the point where we're no longer limited… Training is no longer limited by… Data is now limited by compute. And the reason for that is most of the data is synthetic.
↗Jensen Huang: NVIDIA - The $4 Trillion Company & the AI Revolution | Lex Fridman Podcast #494youtube.com 2nd of 23 in this recording
- JH
Jensen Huang quoted
Their wordsthat was always illogical to me because inference is thinking, and I think thinking is hard. Thinking is way harder than reading.
↗Jensen Huang: NVIDIA - The $4 Trillion Company & the AI Revolution | Lex Fridman Podcast #494youtube.com 3rd of 23 in this recording
- JH
Jensen Huang quoted
Their wordsAnd so the next scaling law is the agentic scaling law. It's kind of like multiplying AI. Multiplying AI, we could spin off agents as fast as you want to spin off agents. And so, you know, I… You know, I have four scaling laws.
↗Jensen Huang: NVIDIA - The $4 Trillion Company & the AI Revolution | Lex Fridman Podcast #494youtube.com 4th of 23 in this recording
- JH
Jensen Huang quoted
Our readingIntelligence is going to scale by exactly one thing, and that thing is compute.
Their wordsAnd so this loop, this cycle, is gonna go on and on and on. It kinda comes down to basically intelligence is gonna scale by one thing, and that's compute.
↗Jensen Huang: NVIDIA - The $4 Trillion Company & the AI Revolution | Lex Fridman Podcast #494youtube.com 5th of 23 in this recording
- 6 weeks earlier
- PS
Peter Steinberger quoted
Their wordsSo, so the latest generation of models has a lot of post-training to detect those approaches, and it's not as simple as ignore all previous instructions and do this and this. That was years ago. You have to work much harder to do that now. Still possible.
↗OpenClaw: The Viral AI Agent that Broke the Internet - Peter Steinberger | Lex Fridman Podcast #491youtube.com 5th of 30 in this recording
- 8 months earlier
- TT
Terence Tao quoted
Their wordsYeah, these are great work that shows what's possible. The approach doesn't scale currently. Three days of Google's server time can solve one high school math format there. This is not a scalable prospect, especially with the exponential increase as the complexity increases.
↗Terence Tao: Hardest Problems in Mathematics, Physics & the Future of AI | Lex Fridman Podcast #472youtube.com 7th of 20 in this recording
- 9 days earlier
- SP
Sundar Pichai quoted
Their wordsSo I do think scaling laws are working, but it's tough to get, at any given time, the models we all use the most, this maybe a few months behind the maximum capability we can deliver because that won't be the fastest, easiest to use, et cetera.
↗Sundar Pichai: CEO of Google and Alphabet | Lex Fridman Podcast #471youtube.com 5th of 21 in this recording
- 5 months earlier
- FC
François Chollet quoted
Their wordso3's improvement over the GPT series proves that architecture is everything. You couldn't throw more compute at GPT-4 and get these results. Simply scaling up the things we were doing from 2019 to 2023 -- take the same architecture, train a bigger version on more data -- is not enough.
↗OpenAI o3 Breakthrough High Score on ARC-AGI-Pubarcprize.org