Tap a claim on the ring to put it at the centre.
← Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
17 connected korrents · 13 moments on record from 12 Jun 2017 to 12 Aug 2026.
Everything filed under neural networks
neural networks
Everything filed under LLMs
LLMs
Everything filed under measuring intelligence
measuring intelligence
Everything filed under reinforcement learning
reinforcement learning
Everything filed under robotics
robotics
Everything filed under transformers
transformers
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 6 years ago
28 May 2020
DA
Dario Amodei — holds since 2020-05-28 — tap for who they are
Same subject: The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out. — tap to centre the map on it
The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: It is the diversity of a robot data set, not its size, that produces generalization: dropping the most diverse slice hurts far more than dropping a random fifth. — tap to centre the map on it
It is the diversity of a robot data set, not its size, that produces generalization: dropping the most diverse slice hurts far more than dropping a random fifth.
Last stated 4 weeks ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from. — tap to centre the map on it
Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning. — tap to centre the map on it
We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning.
Last stated 11 months ago
26 Sept 2025
RS
Richard Sutton — holds since 2025-09-26 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: AI does not remove a job so much as change what it consists of, and knowing exactly where a model's capabilities stop is now part of being a good engineer. — tap to centre the map on it
AI does not remove a job so much as change what it consists of, and knowing exactly where a model's capabilities stop is now part of being a good engineer.
Last stated 11 months ago
16 Oct 2025
DF
Dylan Field — holds since 2025-10-16 — tap for who they are
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel. — tap to centre the map on it
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model. — tap to centre the map on it
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated a month ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
Same subject: Access to large-scale generalist AI systems that could be weaponized should be limited, meaning their code and neural network parameters should not be released open-source. — tap to centre the map on it
Access to large-scale generalist AI systems that could be weaponized should be limited, meaning their code and neural network parameters should not be released open-source.
Last stated 3 years ago
24 Jun 2023
YB
Yoshua Bengio — holds since 2023-06-24 — tap for who they are
Same subject: An agent can understand a maximally simplified explanation and still be unable to come up with it — that gap is what is left of the expert's job. — tap to centre the map on it
An agent can understand a maximally simplified explanation and still be unable to come up with it — that gap is what is left of the expert's job.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: Anything nature shaped can be learned efficiently by a classical neural network, because evolutionary processes leave structure behind. — tap to centre the map on it
Anything nature shaped can be learned efficiently by a classical neural network, because evolutionary processes leave structure behind.
Last stated a year ago
23 Jul 2025
DH
Demis Hassabis — holds since 2025-07-23 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Read this korrent →
Similar wording
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 11 Oct 2018 · 8 years ago
Holds KT Kristina ToutanovaMC Ming-Wei ChangKL Kenton LeeJD Jacob Devlin
Similar wording
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 28 May 2020 · 6 years ago
Holds DA Dario Amodei
Similar wording
The knowledge a model soaks up in pre-training is holding it back; what we actually want is the intelligence with the knowledge stripped out.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Similar wording
It is the diversity of a robot data set, not its size, that produces generalization: dropping the most diverse slice hurts far more than dropping a random fifth.
Last stated 12 Aug 2026 · 4 weeks ago
Holds CF Chelsea Finn
Similar wording
Computer-use agents had to wait for language models: without pre-trained representations the reward is too sparse to ever learn from.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Similar wording
We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning.
Last stated 26 Sept 2025 · 11 months ago
Holds RS Richard Sutton
Similar wording
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 12 Jun 2017 · 9 years ago
Holds IP Illia PolosukhinJU Jakob UszkoreitNP Niki ParmarAV Ashish VaswaniNS Noam ShazeerAG Aidan N. GomezLJ Llion JonesŁK Łukasz Kaiser
Similar wording
AI does not remove a job so much as change what it consists of, and knowing exactly where a model's capabilities stop is now part of being a good engineer.
Last stated 16 Oct 2025 · 11 months ago
Holds DF Dylan Field
Same subject: measuring intelligence
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: measuring intelligence
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 5 Nov 2019 · 7 years ago
Holds François Chollet
Same subject: reinforcement learning
A physical AI company needs three AIs, not one — the agent, the simulator and the critic — turning deployment into a flywheel.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Building a realistic simulator is exactly as hard as building the agent, because the simulator is itself a large AI model.
Last stated 3 Aug 2026 · a month ago
Holds DD Dmitri Dolgov
Same subject: neural networks
Access to large-scale generalist AI systems that could be weaponized should be limited, meaning their code and neural network parameters should not be released open-source.
Last stated 24 Jun 2023 · 3 years ago
Holds Yoshua Bengio
Same subject: neural networks
An agent can understand a maximally simplified explanation and still be unable to come up with it — that gap is what is left of the expert's job.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: neural networks
Anything nature shaped can be learned efficiently by a classical neural network, because evolutionary processes leave structure behind.
Last stated 23 Jul 2025 · a year ago
Holds DH Demis Hassabis