korrents

On the map

Tap a claim on the ring to put it at the centre.

← A transduction model can rely entirely on self-attention for input and…

8 connected korrents · 4 moments on record from 12 Jun 2017 to 15 Jul 2026.

Everything filed under transformers transformers Everything filed under LLMs LLMs Everything filed under AI alignment AI alignment Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. A transduction model can rely entirely onself-attention for input and outputrepresentations without sequence-alignedRNNs or convolution. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are LJ Llion Jones — holds since 2017-06-12 — tap for who they are ŁK Łukasz Kaiser — holds since 2017-06-12 — tap for who they are Same subject: Context engineering has stayed relevant for a year because it is grounded in how transformer attention works, and it will matter to anyone building on AI until post-transformer models arrive. — tap to centre the map on it Context engineering has stayedrelevant for a year because it isgrounded in how transformerattention works, and it will matterto anyone building on AI untilpost-transformer models arrive. Last stated 2 months ago 15 Jul 2026 DH Dex Horthy — holds since 2026-07-15 — tap for who they are Same subject: Self-attention can yield more interpretable models. — tap to centre the map on it Self-attention can yield moreinterpretable models. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Sequence transduction can use a simple architecture based solely on attention, without recurrence or convolutions. — tap to centre the map on it Sequence transduction can use asimple architecture based solely onattention, without recurrence orconvolutions. Last stated 9 years ago 12 Jun 2017 IP Illia Polosukhin — holds since 2017-06-12 — tap for who they are JU Jakob Uszkoreit — holds since 2017-06-12 — tap for who they are NP Niki Parmar — holds since 2017-06-12 — tap for who they are AV Ashish Vaswani — holds since 2017-06-12 — tap for who they are NS Noam Shazeer — holds since 2017-06-12 — tap for who they are AG Aidan N. Gomez — holds since 2017-06-12 — tap for who they are +2 2 more on record Same subject: Current techniques restrict pre-trained representation power because standard language models are unidirectional. — tap to centre the map on it Current techniques restrictpre-trained representation powerbecause standard language models areunidirectional. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: Pre-trained representations reduce the need for many heavily-engineered task-specific architectures. — tap to centre the map on it Pre-trained representations reducethe need for many heavily-engineeredtask-specific architectures. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: A deep bidirectional model is strictly more powerful than a left-to-right model or a shallow concatenation of unidirectional models. — tap to centre the map on it A deep bidirectional model isstrictly more powerful than aleft-to-right model or a shallowconcatenation of unidirectionalmodels. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are Same subject: The next-token language modeling objective is misaligned with following user instructions helpfully and safely. — tap to centre the map on it The next-token language modelingobjective is misaligned withfollowing user instructionshelpfully and safely. Last stated 5 years ago 4 Mar 2022 JL Jan Leike — holds since 2022-03-04 — tap for who they are Same subject: Unidirectional restrictions are sub-optimal for sentence-level tasks and harmful for token-level tasks that need bidirectional context. — tap to centre the map on it Unidirectional restrictions aresub-optimal for sentence-level tasksand harmful for token-level tasksthat need bidirectional context. Last stated 8 years ago 11 Oct 2018 KT Kristina Toutanova — holds since 2018-10-11 — tap for who they are MC Ming-Wei Chang — holds since 2018-10-11 — tap for who they are KL Kenton Lee — holds since 2018-10-11 — tap for who they are JD Jacob Devlin — holds since 2018-10-11 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution. Last stated 12 Jun 2017 · 9 years ago Holds Illia PolosukhinJakob UszkoreitNiki ParmarAshish VaswaniNoam ShazeerAidan N. GomezLlion JonesŁukasz Kaiser Read this korrent →