Tap a claim on the ring to put it at the centre.
← Self-attention can yield more interpretable models.
9 connected korrents · 8 moments on record from 12 Jun 2017 to 12 Aug 2026.
Everything filed under transformers
transformers
Everything filed under LLMs
LLMs
Everything filed under AI and human skill
AI and human skill
Everything filed under robotics
robotics
Everything filed under recursive self-improvement
recursive self-improvement
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Self-attention can yield more interpretable models.
Self-attention can yield more interpretable models.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
LJ
Llion Jones — holds since 2017-06-12 — tap for who they are
ŁK
Łukasz Kaiser — holds since 2017-06-12 — tap for who they are
Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Last stated 9 years ago
12 Jun 2017
IP
Illia Polosukhin — holds since 2017-06-12 — tap for who they are
JU
Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are
NP
Niki Parmar — holds since 2017-06-12 — tap for who they are
AV
Ashish Vaswani — holds since 2017-06-12 — tap for who they are
NS
Noam Shazeer — holds since 2017-06-12 — tap for who they are
AG
Aidan N. Gomez — holds since 2017-06-12 — tap for who they are
+2
2 more on record
Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it
Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive.
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: Renouncing your own cognitive ability because the models are good is premature, because companies still pay an enormous premium for it. — tap to centre the map on it
Renouncing your own cognitive ability because the models are good is premature, because companies still pay an enormous premium for it.
Last stated a month ago
31 Jul 2026
PC
Patrick Collison — holds since 2026-07-31 — tap for who they are
Same subject: Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it. — tap to centre the map on it
Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it.
Last stated 4 weeks ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it
A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models.
Last stated 8 years ago
11 Oct 2018
KT
Kristina Toutanova — holds since 2018-10-11 — tap for who they are
MC
Ming-Wei Chang — holds since 2018-10-11 — tap for who they are
KL
Kenton Lee — holds since 2018-10-11 — tap for who they are
JD
Jacob Devlin — holds since 2018-10-11 — tap for who they are
Same subject: There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate. — tap to centre the map on it
There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate.
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: A bigger context window does not give you a smarter model; the intelligence of the model is what decides how much of that window it can actually attend to. — tap to centre the map on it
A bigger context window does not give you a smarter model; the intelligence of the model is what decides how much of that window it can actually attend to.
Last stated 2 months ago
15 Jul 2026
DH
Dex Horthy — holds since 2026-07-15 — tap for who they are
Same subject: A child who has seen ten cats learns what a machine needs the whole internet of cat photos for, by a learning pathway nobody has solved. — tap to centre the map on it
A child who has seen ten cats learns what a machine needs the whole internet of cat photos for, by a learning pathway nobody has solved.
Last stated a month ago
10 Aug 2026
FL
Fei-Fei Li — holds since 2026-08-10 — tap for who they are
Same subject: A frontier model is measurably more intelligent with no system prompt at all; the prompts that remain are there for the product, not the model. — tap to centre the map on it
A frontier model is measurably more intelligent with no system prompt at all; the prompts that remain are there for the product, not the model.
Last stated a month ago
27 Jul 2026
BC
Boris Cherny — holds since 2026-07-27 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are