Tap a claim on the ring to put it at the centre.
← A deep bidirectional model is strictly more powerful than a…
17 connected korrents · 12 moments on record from 10 Jun 2014 to 11 Aug 2026.
Everything filed under scaling laws
scaling laws
Everything filed under measuring intelligence
measuring intelligence
Everything filed under reinforcement learning
reinforcement learning
Everything filed under LLMs
LLMs
Everything filed under transformers
transformers
Everything filed under neural networks
neural networks
Everything filed under mathematics
mathematics
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Unidirectional restrictions are sub-optimal for sentence-level tasks and harmful for token-level tasks that need bidirectional context. — tap to centre the map on it
Unidirectional restrictions are sub-optimal for sentence-level tasks and harmful for token-level tasks that need bidirectional context.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Depth is not free: past a point, deeper neural networks are harder to train rather than simply better. — tap to centre the map on it
Depth is not free: past a point, deeper neural networks are harder to train rather than simply better.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes. — tap to centre the map on it
Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: What sits in a model's context is far clearer to it than its training data, which is trillions of tokens stirred into a soup of parameters. — tap to centre the map on it
What sits in a model's context is far clearer to it than its training data, which is trillions of tokens stirred into a soup of parameters.
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: Deep generative models have had less impact than discriminative models because of intractable probabilistic computations and difficulty using piecewise linear units. — tap to centre the map on it
Deep generative models have had less impact than discriminative models because of intractable probabilistic computations and difficulty using piecewise linear units.
Last stated 12 years ago
10 Jun 2014
YB
Yoshua Bengio — holds since 2014-06-10 — tap for who they are
Same subject: Large language models start in exactly the wrong place, because they try to get by without a goal and therefore without any sense of better or worse. — tap to centre the map on it
Large language models start in exactly the wrong place, because they try to get by without a goal and therefore without any sense of better or worse.
Last stated 11 months ago
26 Sept 2025
RS
Richard Sutton — holds since 2025-09-26 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget. — tap to centre the map on it
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Read this korrent →
Similar wording
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Similar wording
Unidirectional restrictions are sub-optimal for sentence-level tasks and harmful for token-level tasks that need bidirectional context.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Similar wording
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 12 Jun 2017 · 9 years ago
Holds IP Illia PolosukhinJU Jakob UszkoreitNP Niki ParmarAV Ashish VaswaniNS Noam ShazeerAG Aidan N. GomezLJ Llion JonesŁK Łukasz Kaiser
Similar wording
Depth is not free: past a point, deeper neural networks are harder to train rather than simply better.
Last stated 10 Dec 2015 · 11 years ago
Holds JS Jian SunKH Kaiming He
Similar wording
Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
What sits in a model's context is far clearer to it than its training data, which is trillions of tokens stirred into a soup of parameters.
Last stated 30 Jul 2026 · a month ago
Holds JD Jeff Dean
Similar wording
Deep generative models have had less impact than discriminative models because of intractable probabilistic computations and difficulty using piecewise linear units.
Last stated 10 Jun 2014 · 12 years ago
Holds Yoshua Bengio
Similar wording
Large language models start in exactly the wrong place, because they try to get by without a goal and therefore without any sense of better or worse.
Last stated 26 Sept 2025 · 11 months ago
Holds RS Richard Sutton
Same subject: scaling laws
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: scaling laws
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 13 Mar 2026 · 6 months ago
Holds DP Dylan Patel
Same subject: scaling laws
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: measuring intelligence
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: reinforcement learning
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov