Tap a claim on the ring to put it at the centre.
← Looped/recurrent-depth transformer architectures can improve model…
17 connected korrents · 17 moments on record from 10 Dec 2015 to 15 Sept 2026.
Everything filed under scaling laws
scaling laws
Everything filed under measuring intelligence
measuring intelligence
Everything filed under reinforcement learning
reinforcement learning
Everything filed under LLMs
LLMs
Everything filed under neural networks
neural networks
Everything filed under transformers
transformers
Everything filed under recursive self-improvement
recursive self-improvement
Everything filed under robotics
robotics
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough.
Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough.
Last stated 2 weeks ago
9 Sept 2026
SR
Sebastian Raschka — holds since 2026-09-09 — tap for who they are
Same subject: Test-time scaling has gained a third axis of latent space reasoning iterations in looped transformers. — tap to centre the map on it
Test-time scaling has gained a third axis of latent space reasoning iterations in looped transformers.
Last stated 2 weeks ago
5 Sept 2026
FC
François Chollet — holds since 2026-09-05 — tap for who they are
Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 6 years ago
28 May 2020
DA
Dario Amodei — holds since 2020-05-28 — tap for who they are
Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it
Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them. — tap to centre the map on it
Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them.
Last stated 4 months ago
16 May 2026
SR
Sebastian Raschka — holds since 2026-05-16 — tap for who they are
Same subject: As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary — tap to centre the map on it
As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary
Last stated 3 years ago
16 Jan 2024
CH
Chip Huyen — holds since 2024-01-16 — tap for who they are
Same subject: There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate. — tap to centre the map on it
There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate.
Last stated 2 months ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it. — tap to centre the map on it
Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it.
Last stated a month ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: "Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next. — tap to centre the map on it
"Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next.
Last stated 10 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 3 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it
A true artificial general intelligence cannot exist without being recognized as a moral subject.
Last stated a year ago
10 Jun 2025
SH
Samuel Hammond — holds since 2025-06-10 — tap for who they are
Same subject: A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model. — tap to centre the map on it
A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model.
Last stated 6 days ago
15 Sept 2026
PG
Paul Graham — holds since 2026-09-15 — tap for who they are
Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it
Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 2 months ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 2 months ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Looped/recurrent-depth transformer architectures can improve model quality at a fixed compute budget when the model is large enough.
Last stated 9 Sept 2026 · 2 weeks ago
Holds Sebastian Raschka
Read this korrent →
Similar wording
Test-time scaling has gained a third axis of latent space reasoning iterations in looped transformers.
Last stated 5 Sept 2026 · 2 weeks ago
Holds François Chollet
Similar wording
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 28 May 2020 · 6 years ago
Holds Dario Amodei
Similar wording
Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases.
Last stated 10 Dec 2015 · 11 years ago
Holds Jian Sun Kaiming He
Similar wording
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 12 Jun 2017 · 9 years ago
Holds Illia Polosukhin Jakob Uszkoreit Niki Parmar Ashish Vaswani Noam Shazeer Aidan N. Gomez Llion Jones Łukasz Kaiser
Similar wording
Larger language models already have enough capacity that extra per-layer embedding tricks add little benefit to them.
Last stated 16 May 2026 · 4 months ago
Holds Sebastian Raschka
Similar wording
As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary
Last stated 16 Jan 2024 · 3 years ago
Holds Chip Huyen
Similar wording
There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate.
Last stated 30 Jul 2026 · 2 months ago
Holds Jeff Dean
Similar wording
Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it.
Last stated 12 Aug 2026 · a month ago
Holds Chelsea Finn
Same subject: scaling laws
"Scaling" was powerful because it was one word: naming a research direction is what tells a whole field what to do next.
Last stated 25 Nov 2025 · 10 months ago
Holds Ilya Sutskever
Same subject: scaling laws
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 3 months ago
Holds Noam Brown
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds Dylan Patel
Same subject: measuring intelligence
A true artificial general intelligence cannot exist without being recognized as a moral subject.
Last stated 10 Jun 2025 · a year ago
Holds Samuel Hammond
Same subject: measuring intelligence
A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model.
Last stated 15 Sept 2026 · 6 days ago
Holds Paul Graham
Same subject: measuring intelligence
Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: reinforcement learning
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 3 Aug 2026 · 2 months ago
Holds Dmitri Dolgov
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 3 Aug 2026 · 2 months ago
Holds Dmitri Dolgov