transformers
The architecture that dropped recurrence in favour of attention, and that almost every large model since is built on.
FilterEveryone, all time
- DH
Dex Horthy quoted
Their wordsI think context engineering has been so long lived because it's it's grounded in the fundamentals of how transformer attention works and until we have post transformer models or linear attention or whatever it is which who knows when that's going to happen context engineering will be interesting and important to anyone building on AI
↗Context engineering with Dex Horthyyoutube.com 3rd of 26 in this recording
- 9 years earlier
- AV
Ashish Vaswani quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- NS
Noam Shazeer quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- NP
Niki Parmar quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- JU
Jakob Uszkoreit quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- LJ
Llion Jones quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- AG
Aidan N. Gomez quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- ŁK
Łukasz Kaiser quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- IP
Illia Polosukhin quoted
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
- AV
Ashish Vaswani quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
- NS
Noam Shazeer quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 10th of 24 in this piece
- NP
Niki Parmar quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 11th of 24 in this piece
- JU
Jakob Uszkoreit quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 12th of 24 in this piece
- LJ
Llion Jones quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 13th of 24 in this piece
- AG
Aidan N. Gomez quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 14th of 24 in this piece
- ŁK
Łukasz Kaiser quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 15th of 24 in this piece
- IP
Illia Polosukhin quoted
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 16th of 24 in this piece
- AV
Ashish Vaswani quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 17th of 24 in this piece
- NS
Noam Shazeer quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 18th of 24 in this piece
- NP
Niki Parmar quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 19th of 24 in this piece
- JU
Jakob Uszkoreit quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 20th of 24 in this piece
- LJ
Llion Jones quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 21st of 24 in this piece
- AG
Aidan N. Gomez quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 22nd of 24 in this piece
- ŁK
Łukasz Kaiser quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 23rd of 24 in this piece
- IP
Illia Polosukhin quoted
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 24th of 24 in this piece