Tap a claim on the ring to put it at the centre.
← With suitable architecture, gradient-based learning can form decision…
15 connected korrents · 11 moments on record from 1 Jan 1998 to 26 Jun 2026.
Everything filed under measuring intelligence
measuring intelligence
Everything filed under pruning
pruning
Everything filed under scaling laws
scaling laws
Everything filed under neural networks
neural networks
Everything filed under optimizers
optimizers
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: With suitable architecture, gradient-based learning can form decision surfaces that classify high-dimensional patterns such as handwritten characters with minimal preprocessing.
With suitable architecture, gradient-based learning can form decision surfaces that classify high-dimensional patterns such as handwritten characters with minimal preprocessing.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: Multilayer neural networks trained with backpropagation are the best example of successful gradient-based learning. — tap to centre the map on it
Multilayer neural networks trained with backpropagation are the best example of successful gradient-based learning.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures. — tap to centre the map on it
Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Gradient descent finds a solution to the problems a model has seen; nothing in the algorithm makes it pick the one that generalises well. — tap to centre the map on it
Gradient descent finds a solution to the problems a model has seen; nothing in the algorithm makes it pick the one that generalises well.
Last stated 11 months ago
26 Sept 2025
RS
Richard Sutton — holds since 2025-09-26 — tap for who they are
Same subject: Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance. — tap to centre the map on it
Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train. — tap to centre the map on it
An optimizer that is invariant to rescaling the gradients and cheap in memory is what makes very large models practical to train.
Last stated 12 years ago
22 Dec 2014
JB
Jimmy Ba — holds since 2014-12-22 — tap for who they are
DK
Diederik P. Kingma — holds since 2014-12-22 — tap for who they are
Same subject: Convolutional neural networks are specifically designed to handle variability in two-dimensional shapes. — tap to centre the map on it
Convolutional neural networks are specifically designed to handle variability in two-dimensional shapes.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence. — tap to centre the map on it
Language models are already a form of AGI; what the labs are chasing is a further step, not the arrival of general intelligence.
Last stated 2 years ago
3 Feb 2025
NL
Nathan Lambert — holds since 2025-02-03 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it. — tap to centre the map on it
A small enough winning ticket learns faster than the network it was cut out of, and ends up more accurate than it.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work. — tap to centre the map on it
A winning ticket wins on its initial weights: the connections it keeps started at values that happen to make training work.
Last stated 9 years ago
9 Mar 2018
MC
Michael Carbin — holds since 2018-03-09 — tap for who they are
JF
Jonathan Frankle — holds since 2018-03-09 — tap for who they are
Same subject: Batteries will replace transmission as the cheapest way to keep the lights on, and the grid will shrink rather than grow. — tap to centre the map on it
Batteries will replace transmission as the cheapest way to keep the lights on, and the grid will shrink rather than grow.
Last stated 9 months ago
8 Dec 2025
CH
Casey Handmer — holds since 2023-10-11 — tap for who they are
CH
Casey Handmer — holds since 2025-12-08 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from. — tap to centre the map on it
A lab should spend most of its compute on research rather than on building the next model, because research is where the tenfold yearly efficiency gains come from.
Last stated 6 months ago
13 Mar 2026
DP
Dylan Patel — holds since 2026-03-13 — tap for who they are
Same subject: A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget. — tap to centre the map on it
A model's capability is now a function of how much money you spend on it, so asking what a model can do means nothing until you name a budget.
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 1998 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are