Tap a claim on the ring to put it at the centre.
← The gains from extra simulated layers in AI are much smaller than the gains from extra real layers.
17 connected korrents · 17 moments from 10 Dec 2015 to 24 Sept 2026.
Everything filed under scaling laws
scaling laws
Everything filed under neural networks
neural networks
Everything filed under LLMs
LLMs
Everything filed under design
design
Everything filed under AI and science
AI and science
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: The gains from extra simulated layers in AI are much smaller than the gains from extra real layers.
The gains from extra simulated layers in AI are much smaller than the gains from extra real layers.
Last stated a week ago
24 Sept 2026
SA
Scott Alexander — holds since 2026-09-24 — tap for who they are
Same subject: AI is grown rather than designed, which is why its overall behaviour resists any description we can fully understand. — tap to centre the map on it
AI is grown rather than designed, which is why its overall behaviour resists any description we can fully understand.
Last stated 4 weeks ago
6 Sept 2026
JP
Jakub Pachocki — holds since 2026-09-06 — tap for who they are
Same subject: A small performance gain is not worth the complexity added to achieve it, since complexity's costs compound rather than add. — tap to centre the map on it
A small performance gain is not worth the complexity added to achieve it, since complexity's costs compound rather than add.
Last stated 3 years ago
6 Dec 2023
RP
Rob Pike — holds since 2023-12-06 — tap for who they are
Same subject: Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them. — tap to centre the map on it
Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them.
Last stated 5 months ago
16 May 2026
SR
Sebastian Raschka — holds since 2026-05-16 — tap for who they are
Same subject: Getting genuinely original output from AI requires working hard against the grain of its defaults. — tap to centre the map on it
Getting genuinely original output from AI requires working hard against the grain of its defaults.
Last stated 3 months ago
29 Jun 2026
KC
Kyle Chayka — holds since 2026-06-29 — tap for who they are
Same subject: The capability gap between smaller, locally-run AI models and larger models will narrow over time as compute becomes more accessible. — tap to centre the map on it
The capability gap between smaller, locally-run AI models and larger models will narrow over time as compute becomes more accessible.
Last stated 5 months ago
27 Apr 2026
JS
Jon Seager — holds since 2026-04-27 — tap for who they are
Same subject: Current AI models still fall short of expert human performance. — tap to centre the map on it
Current AI models still fall short of expert human performance.
Last stated 2 years ago
27 Jan 2025
NE
Nelson Elhage — holds since 2025-01-27 — tap for who they are
Same subject: A deliberately slower pace of AI model releases would still feel faster than the pace of AI improvement that came before it. — tap to centre the map on it
A deliberately slower pace of AI model releases would still feel faster than the pace of AI improvement that came before it.
Last stated 3 weeks ago
14 Sept 2026
BH
Byrne Hobart — holds since 2026-09-14 — tap for who they are
Same subject: AI models can be more accurate than physical experiments by averaging out noise, incorporating priors from large datasets, and interpolating between data points. — tap to centre the map on it
AI models can be more accurate than physical experiments by averaging out noise, incorporating priors from large datasets, and interpolating between data points.
Last stated 10 months ago
23 Nov 2025
EH
Elliot Hershberg — holds since 2025-11-23 — tap for who they are
Same subject: Scaling the 2019-2023 recipe — same architecture, bigger model, more data — is not enough; further progress depends on new architectural ideas. — tap to centre the map on it
Scaling the 2019-2023 recipe — same architecture, bigger model, more data — is not enough; further progress depends on new architectural ideas.
Last stated 2 years ago
20 Dec 2024
FC
François Chollet — holds since 2024-12-20 — tap for who they are
GM
Gary Marcus — holds since 2024-10-11 — tap for who they are
Same subject: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. — tap to centre the map on it
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it. — tap to centre the map on it
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it
"Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next.
Last stated 10 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 7 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work. — tap to centre the map on it
A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: Intelligence is the execution of small, simplicity-weighted programs found by search over smooth loss landscapes in overparameterized networks, approaching Bayes-optimal reasoning as scale increases. — tap to centre the map on it
Intelligence is the execution of small, simplicity-weighted programs found by search over smooth loss landscapes in overparameterized networks, approaching Bayes-optimal reasoning as scale increases.
Last stated 5 years ago
11 Jun 2021
GW
Gwern — holds since 2021-06-11 — tap for who they are
Same subject: Pruning cuts the parameters of a trained network by over 90% without costing accuracy, so inference gets smaller and faster for free. — tap to centre the map on it
Pruning cuts the parameters of a trained network by over 90% without costing accuracy, so inference gets smaller and faster for free.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
The gains from extra simulated layers in AI are much smaller than the gains from extra real layers.
Last stated 24 Sept 2026 · a week ago
Holds Scott Alexander
Read this korrent →
Similar wording
AI is grown rather than designed, which is why its overall behaviour resists any description we can fully understand.
Last stated 6 Sept 2026 · 4 weeks ago
Holds Jakub Pachocki
Similar wording
A small performance gain is not worth the complexity added to achieve it, since complexity's costs compound rather than add.
Last stated 6 Dec 2023 · 3 years ago
Holds Rob Pike
Similar wording
Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them.
Last stated 16 May 2026 · 5 months ago
Holds Sebastian Raschka
Similar wording
Getting genuinely original output from AI requires working hard against the grain of its defaults.
Last stated 29 Jun 2026 · 3 months ago
Holds Kyle Chayka
Similar wording
The capability gap between smaller, locally-run AI models and larger models will narrow over time as compute becomes more accessible.
Last stated 27 Apr 2026 · 5 months ago
Holds Jon Seager
Similar wording
Current AI models still fall short of expert human performance.
Last stated 27 Jan 2025 · 2 years ago
Holds Nelson Elhage
Similar wording
A deliberately slower pace of AI model releases would still feel faster than the pace of AI improvement that came before it.
Last stated 14 Sept 2026 · 3 weeks ago
Holds Byrne Hobart
Similar wording
AI models can be more accurate than physical experiments by averaging out noise, incorporating priors from large datasets, and interpolating between data points.
Last stated 23 Nov 2025 · 10 months ago
Holds Elliot Hershberg
Same subject: neural networks
Scaling the 2019-2023 recipe — same architecture, bigger model, more data — is not enough; further progress depends on new architectural ideas.
Last stated 20 Dec 2024 · 2 years ago
Holds François Chollet Gary Marcus
Same subject: neural networks
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 10 Dec 2015 · 11 years ago
Holds Jian Sun Kaiming He
Same subject: neural networks
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 Mar 2018 · 9 years ago
Holds Michael Carbin Jonathan Frankle
Same subject: scaling laws
"Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next.
Last stated 25 Nov 2025 · 10 months ago
Holds Ilya Sutskever
Same subject: scaling laws
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 7 months ago
Holds Dylan Patel
Same subject: pruning
A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work.
Last stated 9 Mar 2018 · 9 years ago
Holds Michael Carbin Jonathan Frankle
Same subject: pruning
Intelligence is the execution of small, simplicity-weighted programs found by search over smooth loss landscapes in overparameterized networks, approaching Bayes-optimal reasoning as scale increases.
Last stated 11 Jun 2021 · 5 years ago
Holds Gwern
Same subject: pruning
Pruning cuts the parameters of a trained network by over 90% without costing accuracy, so inference gets smaller and faster for free.
Last stated 9 Mar 2018 · 9 years ago
Holds Michael Carbin Jonathan Frankle