korrents

On the map

Tap a claim on the ring to put it at the centre.

← Self-attention can yield more interpretable models.

9 connected korrents · 8 moments on record from 12 Jun 2017 to 12 Aug 2026.

Everything filed under transformers transformers Everything filed under LLMs LLMs Everything filed under AI and human skill AI and human skill Everything filed under robotics robotics Everything filed under recursive self-improvement recursive self-improvement Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Self-attention can yield more interpretable models. Self-attention can yield moreinterpretable models. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are LJ Llion Jones — holds since 2017-06-12 — tap for who they are ŁK Łukasz Kaiser — holds since 2017-06-12 — tap for who they are Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it A transduction model can relyentirely on self-attention for inputand output representations withoutsequence-aligned RNNs orconvolution. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it Context engineering has stayedrelevant for a year because it isgrounded in how transformerattention works, and it will matterto anyone building on AI untilpost-transformer models arrive. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: Renouncing your own cognitive ability because the models are good is premature, because companies still pay an enormous premium for it. — tap to centre the map on it Renouncing your own cognitiveability because the models are goodis premature, because companiesstill pay an enormous premium forit. Last stated a month ago 31 Jul 2026 PC Patrick Collison — holds since 2026-07-31 — tap for who they are Same subject: Letting a robot imagine the next image helps, but it is not essential — the model was surprisingly good without it. — tap to centre the map on it Letting a robot imagine the nextimage helps, but it is not essential— the model was surprisingly goodwithout it. Last stated 4 weeks ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it A deep bidirectional model isstrictly more powerful than aleft-to-right model or a shallowconcatenation of unidirectionalmodels. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: There is no real impediment to models running their own research loop and improving themselves at a far more rapid rate. — tap to centre the map on it There is no real impediment tomodels running their own researchloop and improving themselves at afar more rapid rate. Last stated a month ago 30 Jul 2026 JD Jeff Dean — holds since 2026-07-30 — tap for who they are Same subject: A bigger context window does not give you a smarter model; the intelligence of the model is what decides how much of that window it can actually attend to. — tap to centre the map on it A bigger context window does notgive you a smarter model; theintelligence of the model is whatdecides how much of that window itcan actually attend to. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: A child who has seen ten cats learns what a machine needs the whole internet of cat photos for, by a learning pathway nobody has solved. — tap to centre the map on it A child who has seen ten cats learnswhat a machine needs the wholeinternet of cat photos for, by alearning pathway nobody has solved. Last stated a month ago 10 Aug 2026 FL Fei-Fei Li — holds since 2026-08-10 — tap for who they are Same subject: A frontier model is measurably more intelligent with no system prompt at all; the prompts that remain are there for the product, not the model. — tap to centre the map on it A frontier model is measurably moreintelligent with no system prompt atall; the prompts that remain arethere for the product, not themodel. Last stated a month ago 27 Jul 2026 BC Boris Cherny — holds since 2026-07-27 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Self-attention can yield more interpretable models. Last stated 12 Jun 2017 · 9 years ago Holds Illia PolosukhinJakob UszkoreitNiki ParmarAshish VaswaniNoam ShazeerAidan N. GomezLlion JonesŁukasz Kaiser Read this korrent →