What François Chollet thinks about measuring intelligence
Creator of the Keras deep-learning library and the ARC-AGI benchmark.
Everything they publish, on ppll ↗
François Chollet did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
5 dated positions, 2019 to 2026, in their own words. Our reading of what François Chollet has said — not written or endorsed by them.
-
Our reading · no longer heldThe LLM line of research will reach a capability plateau.
Recalls holdingIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
SaidIn the era of base LLM scaling (2022-2024), I believed the LLM line of research would reach a capability plateau (as later seen with base LLMs). In late 2024, after the o3 test-time compute demo, I changed my views: the new models were showing genuine fluid intelligence, and with this new line of work, the LLM line of research could achieve unbounded capability scaling. "There will be no wall."
- 20 months earlier
-
Their wordsThis "memorize, fetch, apply" paradigm can achieve arbitrary levels of skills at arbitrary tasks given appropriate training data, but it cannot adapt to novelty or pick up new skills on the fly (which is to say that there is no fluid intelligence at play here.)
↗OpenAI o3 Breakthrough High Score on ARC-AGI-Pubarcprize.org
- 5 years earlier
-
Their wordsWe argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to "buy" arbitrary levels of skills for a system, in a way that masks the system's own generalization power.
↗On the Measure of Intelligencearxiv.org 2nd of 5 in this piece
-
Their wordsWe argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans.
↗On the Measure of Intelligencearxiv.org 3rd of 5 in this piece
-
Their wordsIf intelligence lies in the process of acquiring skills, then there is no task X such that skill at X demonstrates intelligence, unless X is actually a meta-task involving skill-acquisition across a broad range of tasks.
↗On the Measure of Intelligencearxiv.org 5th of 5 in this piece