Tap a claim on the ring to put it at the centre.
← Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.
17 connected korrents · 15 moments from 10 Dec 2015 to 9 Sept 2026.
Everything filed under scaling laws
scaling laws
Everything filed under neural networks
neural networks
Everything filed under LLMs
LLMs
Everything filed under OpenAI
OpenAI
Everything filed under AI alignment
AI alignment
Everything filed under transformers
transformers
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.
Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.
Last stated 6 years ago
4 Feb 2021
GW
Gwern — holds since 2021-02-04 — tap for who they are
Same subject: The scaling law for world models was a conviction held in advance, not a discovery; the devils were always in the architecture choices and the data mixture. — tap to centre the map on it
The scaling law for world models was a conviction held in advance, not a discovery; the devils were always in the architecture choices and the data mixture.
Last stated 4 weeks ago
4 Sept 2026
FL
Fei-Fei Li — holds since 2026-09-04 — tap for who they are
Same subject: The scaling curve really is S-shaped, but there are at least two cycles left in it, which means models sixteen times smarter and all knowledge work subsumed. — tap to centre the map on it
The scaling curve really is S-shaped, but there are at least two cycles left in it, which means models sixteen times smarter and all knowledge work subsumed.
Last stated 7 months ago
11 Mar 2026
SY
Steve Yegge — holds since 2026-03-11 — tap for who they are
Same subject: In verifiable domains model capability scaling should remain unbounded, because the space of enumerable patterns is infinite by construction. — tap to centre the map on it
In verifiable domains model capability scaling should remain unbounded, because the space of enumerable patterns is infinite by construction.
Last stated a month ago
27 Aug 2026
FC
François Chollet — holds since 2026-08-27 — tap for who they are
Same subject: Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve. — tap to centre the map on it
Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve.
Last stated a year ago
5 Jun 2025
SP
Sundar Pichai — holds since 2025-06-05 — tap for who they are
Same subject: How far AI systems are scaled must be constrained by how confident we are in their safety. — tap to centre the map on it
How far AI systems are scaled must be constrained by how confident we are in their safety.
Last stated 4 weeks ago
6 Sept 2026
JP
Jakub Pachocki — holds since 2026-09-06 — tap for who they are
Same subject: Full fine-tuning of large language models has become too costly, making parameter-efficient fine-tuning methods necessary. — tap to centre the map on it
Full fine-tuning of large language models has become too costly, making parameter-efficient fine-tuning methods necessary.
Last stated 3 years ago
5 Dec 2023
SR
Sebastian Ruder — holds since 2023-12-05 — tap for who they are
Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 6 years ago
28 May 2020
DA
Dario Amodei — holds since 2020-05-28 — tap for who they are
Same subject: As AI models keep scaling up, the value of any single training source keeps shrinking. — tap to centre the map on it
As AI models keep scaling up, the value of any single training source keeps shrinking.
Last stated 4 weeks ago
7 Sept 2026
KK
Kevin Kelly — holds since 2026-09-07 — tap for who they are
Same subject: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. — tap to centre the map on it
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it. — tap to centre the map on it
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it
Self-attention can yield more interpretable models.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more
Same subject: Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough. — tap to centre the map on it
Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough.
Last stated 3 weeks ago
9 Sept 2026
SR
Sebastian Raschka — holds since 2026-09-09 — tap for who they are
Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 4 weeks ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 3 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone who holds the claim — tap it for who they are
At the centre
Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.
Last stated 4 Feb 2021 · 6 years ago
Holds Gwern
Read this korrent →
Similar wording
The scaling law for world models was a conviction held in advance, not a discovery; the devils were always in the architecture choices and the data mixture.
Last stated 4 Sept 2026 · 4 weeks ago
Holds Fei-Fei Li
Similar wording
The scaling curve really is S-shaped, but there are at least two cycles left in it, which means models sixteen times smarter and all knowledge work subsumed.
Last stated 11 Mar 2026 · 7 months ago
Holds Steve Yegge
Similar wording
In verifiable domains model capability scaling should remain unbounded, because the space of enumerable patterns is infinite by construction.
Last stated 27 Aug 2026 · a month ago
Holds François Chollet
Similar wording
Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve.
Last stated 5 Jun 2025 · a year ago
Holds Sundar Pichai
Similar wording
How far AI systems are scaled must be constrained by how confident we are in their safety.
Last stated 6 Sept 2026 · 4 weeks ago
Holds Jakub Pachocki
Similar wording
Full fine-tuning of large language models has become too costly, making parameter-efficient fine-tuning methods necessary.
Last stated 5 Dec 2023 · 3 years ago
Holds Sebastian Ruder
Similar wording
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 28 May 2020 · 6 years ago
Holds Dario Amodei
Similar wording
As AI models keep scaling up, the value of any single training source keeps shrinking.
Last stated 7 Sept 2026 · 4 weeks ago
Holds Kevin Kelly
Same subject: neural networks
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 10 Dec 2015 · 11 years ago
Holds Jian Sun Kaiming He
Same subject: neural networks
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 Mar 2018 · 9 years ago
Holds Michael Carbin Jonathan Frankle
Same subject: neural networks
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: transformers
Self-attention can yield more interpretable models.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: transformers
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Same subject: transformers
Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough.
Last stated 9 Sept 2026 · 3 weeks ago
Holds Sebastian Raschka
Same subject: OpenAI
A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman
Same subject: OpenAI
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 weeks ago
Holds Zvi Mowshowitz
Same subject: OpenAI
Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers.
Last stated 18 Mar 2024 · 3 years ago
Holds Sam Altman