Tap a claim on the ring to put it at the centre.
← A layer should learn a residual with reference to its own input rather…
17 connected korrents · 13 moments on record from 22 Dec 2014 to 3 Sept 2026.
Everything filed under benchmarks
benchmarks
Everything filed under measuring intelligence
measuring intelligence
Everything filed under neural networks
neural networks
Everything filed under reinforcement learning
reinforcement learning
Everything filed under optimizers
optimizers
Everything filed under LLMs
LLMs
Everything filed under mathematics
mathematics
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases. — tap to centre the map on it
Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. — tap to centre the map on it
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train.
Last stated 12 years ago
22 Dec 2014
JB
Jimmy Ba — holds since 2014-12-22 — tap for who they are
DK
Diederik P. Kingma — holds since 2014-12-22 — tap for who they are
Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out. — tap to centre the map on it
The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: Depth is not free: past a point, deeper neural networks are harder to train rather than simply better. — tap to centre the map on it
Depth is not free: past a point, deeper neural networks are harder to train rather than simply better.
Last stated 11 years ago
10 Dec 2015
JS
Jian Sun — holds since 2015-12-10 — tap for who they are
KH
Kaiming He — holds since 2015-12-10 — tap for who they are
Same subject: Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes. — tap to centre the map on it
Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning. — tap to centre the map on it
We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning.
Last stated 11 months ago
26 Sept 2025
RS
Richard Sutton — holds since 2025-09-26 — tap for who they are
Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence. — tap to centre the map on it
Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence.
Last stated 2 years ago
3 Feb 2025
NL
Nathan Lambert — holds since 2025-02-03 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated a week ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages. — tap to centre the map on it
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated a year ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2014 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
A layer should learn a residual with reference to its own input rather than an unreferenced function, which is what makes great depth trainable.
Last stated 10 Dec 2015 · 11 years ago
Holds JS Jian SunKH Kaiming He
Read this korrent →
Similar wording
Residual networks are easier to optimize than plain ones and keep gaining accuracy as depth increases.
Last stated 10 Dec 2015 · 11 years ago
Holds JS Jian SunKH Kaiming He
Similar wording
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train.
Last stated 22 Dec 2014 · 12 years ago
Holds JB Jimmy BaDK Diederik P. Kingma
Similar wording
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Similar wording
The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Similar wording
Depth is not free: past a point, deeper neural networks are harder to train rather than simply better.
Last stated 10 Dec 2015 · 11 years ago
Holds JS Jian SunKH Kaiming He
Similar wording
Machine learning is a shallow field compared with mathematics: even its most important ideas can be explained in a couple of minutes.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning.
Last stated 26 Sept 2025 · 11 months ago
Holds RS Richard Sutton
Similar wording
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Same subject: reinforcement learning
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Same subject: reinforcement learning
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Same subject: reinforcement learning
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: measuring intelligence
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence.
Last stated 3 Feb 2025 · 2 years ago
Holds NL Nathan Lambert
Same subject: measuring intelligence
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · a week ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: benchmarks
A speed difference between languages that share the LLVM backend measures the benchmark author, not the languages.
Last stated 22 Mar 2025 · a year ago
Holds TH ThePrimeagen