korrents

On the map

Tap a claim on the ring to put it at the centre.

← Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions.

16 connected korrents · 11 moments on record from 1 Jan 1998 to 15 Jul 2026.

Everything filed under Google Google Everything filed under LLMs LLMs Everything filed under measuring intelligence measuring intelligence Everything filed under transformers transformers Everything filed under optimizers optimizers Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. Sequence transduction can use a simplearchitecture based solely on attention,without recurrence or convolutions. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are LJ Llion Jones — holds since 2017-06-12 — tap for who they are ŁK Łukasz Kaiser — holds since 2017-06-12 — tap for who they are Same subject: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. — tap to centre the map on it A transduction model can relyentirely on self-attention for inputand output representations withoutsequence-aligned RNNs orconvolution. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Graph transformer networks allow multimodule document recognition systems to be trained globally with gradient methods to minimize overall performance. — tap to centre the map on it Graph transformer networks allowmultimodule document recognitionsystems to be trained globally withgradient methods to minimize overallperformance. Last stated 29 years ago 1 Jan 1998 YB Yoshua Bengio — holds since 1998-01-01 — tap for who they are Same subject: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures. — tap to centre the map on it Pre-trained representations reducethe need for many heavily-engineeredtask-specific architectures. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: We have almost no automated techniques for making a system transfer what it learns, and none of the few we have are used in modern deep learning. — tap to centre the map on it We have almost no automatedtechniques for making a systemtransfer what it learns, and none ofthe few we have are used in moderndeep learning. Last stated 11 months ago 26 Sept 2025 RS Richard Sutton — holds since 2025-09-26 — tap for who they are Same subject: Scaling up language models greatly improves task-agnostic few-shot performance, sometimes matching prior fine-tuning approaches. — tap to centre the map on it Scaling up language models greatlyimproves task-agnostic few-shotperformance, sometimes matchingprior fine-tuning approaches. Last stated 6 years ago 28 May 2020 DA Dario Amodei — holds since 2020-05-28 — tap for who they are Same subject: Stochastic optimization can be done well using only adaptive estimates of the gradient's lower-order moments. — tap to centre the map on it Stochastic optimization can be donewell using only adaptive estimatesof the gradient's lower-ordermoments. Last stated 12 years ago 22 Dec 2014 JB Jimmy Ba — holds since 2014-12-22 — tap for who they are DK Diederik P. Kingma — holds since 2014-12-22 — tap for who they are Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it A deep bidirectional model isstrictly more powerful than aleft-to-right model or a shallowconcatenation of unidirectionalmodels. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it Self-attention can yield moreinterpretable models. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it Context engineering has stayedrelevant for a year because it isgrounded in how transformerattention works, and it will matterto anyone building on AI untilpost-transformer models arrive. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: If intelligence is the process of acquiring skills, no single task demonstrates intelligence unless it is a meta-task of skill-acquisition across many tasks. — tap to centre the map on it If intelligence is the process ofacquiring skills, no single taskdemonstrates intelligence unless itis a meta-task of skill-acquisitionacross many tasks. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: Solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience. — tap to centre the map on it Solely measuring skill at any giventask falls short of measuringintelligence, because skill isheavily modulated by prior knowledgeand experience. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: The Abstraction and Reasoning Corpus can measure human-like general fluid intelligence and enable fair comparisons between AI systems and humans. — tap to centre the map on it The Abstraction and Reasoning Corpuscan measure human-like general fluidintelligence and enable faircomparisons between AI systems andhumans. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it A company's staff-engineer barshould be set against the bestcompanies in the industry ratherthan against its own history, whichis what makes title inflation a realcost. Last stated 5 months ago 1 Apr 2026 TP Thuan Pham — holds since 2026-04-01 — tap for who they are Same subject: A crewed rocket cannot be made safe by making the booster reliable, so the only real way to improve safety is to carry an escape system. — tap to centre the map on it A crewed rocket cannot be made safeby making the booster reliable, sothe only real way to improve safetyis to carry an escape system. Last stated 3 years ago 14 Dec 2023 JB Jeff Bezos — holds since 2023-12-14 — tap for who they are Same subject: A monopolist that can no longer grow by winning new users can only grow by making its product worse for the users it already has. — tap to centre the map on it A monopolist that can no longer growby winning new users can only growby making its product worse for theusers it already has. Last stated 3 years ago 28 Jul 2023 CD Cory Doctorow — holds since 2023-07-28 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 1998 to today (stretched back to the oldest claim here) — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. Last stated 12 Jun 2017 · 9 years ago Holds Illia PolosukhinJakob UszkoreitNiki ParmarAshish VaswaniNoam ShazeerAidan N. GomezLlion JonesŁukasz Kaiser Read this korrent →