Tap a claim on the ring to put it at the centre.
← Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
16 connected korrents · 11 moments on record from 1 Jan 1998 to 15 Jul 2026.
Everything filed under Google
Google
Everything filed under LLMs
LLMs
Everything filed under measuring intelligence
measuring intelligence
Everything filed under transformers
transformers
Everything filed under optimizers
optimizers
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
LJ
Llion Jones — holds since 2017-06-12 — tap for who they are
ŁK
Łukasz Kaiser — holds since 2017-06-12 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it
Current techniques restrict pre-trained representation power because standard language models are unidirectional.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance. — tap to centre the map on it
Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance.
Last stated 29 years ago
1 Jan 1998
YB
Yoshua Bengio — holds since 1998-01-01 — tap for who they are
Same subject: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures. — tap to centre the map on it
Pre-trained representations reduce the need for many heavily-engineered task-specific architectures.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning. — tap to centre the map on it
We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning.
Last stated 11 months ago
26 Sept 2025
RS
Richard Sutton — holds since 2025-09-26 — tap for who they are
Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it
Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches.
Last stated 6 years ago
28 May 2020
DA
Dario Amodei — holds since 2020-05-28 — tap for who they are
Same subject: Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. — tap to centre the map on it
Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments.
Last stated 12 years ago
22 Dec 2014
JB
Jimmy Ba — holds since 2014-12-22 — tap for who they are
DK
Diederik P. Kingma — holds since 2014-12-22 — tap for who they are
Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it
Self-attention can yield more interpretable models.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it
If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it
Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it
The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans.
Last stated 7 years ago
5 Nov 2019
FC
François Chollet — holds since 2019-11-05 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it
A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system.
Last stated 3 years ago
14 Dec 2023
JB
Jeff Bezos — holds since 2023-12-14 — tap for who they are
Same subject: A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has. — tap to centre the map on it
A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has.
Last stated 3 years ago
28 Jul 2023
CD
Cory Doctorow — holds since 2023-07-28 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 1998 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are