korrents

On the map

Tap a claim on the ring to put it at the centre.

← Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.

17 connected korrents · 15 moments from 10 Dec 2015 to 9 Sept 2026.

Everything filed under scaling laws scaling laws Everything filed under neural networks neural networks Everything filed under LLMs LLMs Everything filed under OpenAI OpenAI Everything filed under AI alignment AI alignment Everything filed under transformers transformers Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further. Mixture-of-experts models can scale stablybut begin to show fundamental limits asthey are scaled further. Last stated 6 years ago 4 Feb 2021 GW Gwern — holds since 2021-02-04 — tap for who they are Same subject: The scaling law for world models was a conviction held in advance, not a discovery; the devils were always in the architecture choices and the data mixture. — tap to centre the map on it The scaling law for world models wasa conviction held in advance, not adiscovery; the devils were always inthe architecture choices and thedata mixture. Last stated 4 weeks ago 4 Sept 2026 FL Fei-Fei Li — holds since 2026-09-04 — tap for who they are Same subject: The scaling curve really is S-shaped, but there are at least two cycles left in it, which means models sixteen times smarter and all knowledge work subsumed. — tap to centre the map on it The scaling curve really isS-shaped, but there are at least twocycles left in it, which meansmodels sixteen times smarter and allknowledge work subsumed. Last stated 7 months ago 11 Mar 2026 SY Steve Yegge — holds since 2026-03-11 — tap for who they are Same subject: In verifiable domains model capability scaling should remain unbounded, because the space of enumerable patterns is infinite by construction. — tap to centre the map on it In verifiable domains modelcapability scaling should remainunbounded, because the space ofenumerable patterns is infinite byconstruction. Last stated a month ago 27 Aug 2026 FC François Chollet — holds since 2026-08-27 — tap for who they are Same subject: Scaling laws are still working, but the models people actually use run a few months behind the maximum capability Google can deliver, because the biggest model is too slow and expensive to serve. — tap to centre the map on it Scaling laws are still working, butthe models people actually use run afew months behind the maximumcapability Google can deliver,because the biggest model is tooslow and expensive to serve. Last stated a year ago 5 Jun 2025 SP Sundar Pichai — holds since 2025-06-05 — tap for who they are Same subject: How far AI systems are scaled must be constrained by how confident we are in their safety. — tap to centre the map on it How far AI systems are scaled mustbe constrained by how confident weare in their safety. Last stated 4 weeks ago 6 Sept 2026 JP Jakub Pachocki — holds since 2026-09-06 — tap for who they are Same subject: Full fine-tuning of large language models has become too costly, making parameter-efficient fine-tuning methods necessary. — tap to centre the map on it Full fine-tuning of large languagemodels has become too costly, makingparameter-efficient fine-tuningmethods necessary. Last stated 3 years ago 5 Dec 2023 SR Sebastian Ruder — holds since 2023-12-05 — tap for who they are Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it Scaling up language models greatlyimproves task-agnostic few-shotperformance, sometimes matchingprior fine-tuning approaches. Last stated 6 years ago 28 May 2020 DA Dario Amodei — holds since 2020-05-28 — tap for who they are Same subject: As AI models keep scaling up, the value of any single training source keeps shrinking. — tap to centre the map on it As AI models keep scaling up, thevalue of any single training sourcekeeps shrinking. Last stated 4 weeks ago 7 Sept 2026 KK Kevin Kelly — holds since 2026-09-07 — tap for who they are Same subject: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable. — tap to centre the map on it A layer should learn a residual withreference to its own input ratherthan an unreferenced function, whichis what makes great depth trainable. Last stated 11 years ago 10 Dec 2015 JS Jian Sun — holds since 2015-12-10 — tap for who they are KH Kaiming He — holds since 2015-12-10 — tap for who they are Same subject: A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it. — tap to centre the map on it A small enough winning ticket learnsfaster than the network it was cutout of, and ends up more accuratethan it. Last stated 9 years ago 9 Mar 2018 MC Michael Carbin — holds since 2018-03-09 — tap for who they are JF Jonathan Frankle — holds since 2018-03-09 — tap for who they are Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it A transduction model can relyentirely on self-attention for inputand output representations withoutsequence-aligned RNNs orconvolution. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it Self-attention can yield moreinterpretable models. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it Sequence transduction can use asimple architecture based solely onattention, without recurrence orconvolutions. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more Same subject: Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough. — tap to centre the map on it Looped/recurrent-depth transformerarchitectures can improve modelquality at a fixed compute budgetwhen the model is large enough. Last stated 3 weeks ago 9 Sept 2026 SR Sebastian Raschka — holds since 2026-09-09 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it A technique that lets an AI model'sreasoning shift outside its visibleChain of Thought is dangerous, bothbecause it works and because aleading lab is willing to deploy it. Last stated 4 weeks ago 3 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone who holds the claim — tap it for who they are

At the centre Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further. Last stated 4 Feb 2021 · 6 years ago Holds Gwern Read this korrent →