What François Chollet thinks about scaling laws
Creator of the Keras deep-learning library and the ARC-AGI benchmark.
Everything they publish, on ppll ↗
François Chollet did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
5 dated positions, 2024 to 2026, in their own words. Our reading of what François Chollet has said — not written or endorsed by them.
-
Their wordsSeems like test time scaling has gained a 3rd axis: latent space reasoning iterations in looped transformers.
- 9 days earlier
-
Their wordsIn verifiable domains, model capability scaling should remain unbounded. Models will simply keep improving by "absorbing more and more of the computational universe", which is infinite by construction.
- 3 weeks earlier
-
Our reading · no longer heldThe LLM line of research will reach a capability plateau.
Recalls holdingIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
SaidIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
-
Their wordsI believe current techniques are 4-6 orders of magnitude away from optimality in terms of data efficiency and test-time compute efficiency. But far future AI will be near-optimal.
- 20 months earlier
-
Their wordso3's improvement over the GPT series proves that architecture is everything. You couldn't throw more compute at GPT-4 and get these results. Simply scaling up the things we were doing from 2019 to 2023 -- take the same architecture, train a bigger version on more data -- is not enough.
↗OpenAI o3 Breakthrough High Score on ARC-AGI-Pubarcprize.org