korrents

transformers

The architecture that dropped recurrence in favour of attention, and that almost every large model since is built on.

What people on korrents have said about transformers, newest first — 25 positions from 9 people.

FilterEveryone, all time
  1. DH

    Dex Horthy quoted

    I think context engineering has been so long lived because it's it's grounded in the fundamentals of how transformer attention works and until we have post transformer models or linear attention or whatever it is which who knows when that's going to happen context engineering will be interesting and important to anyone building on AI

    Context engineering with Dex Horthyyoutube.com 3rd of 26 in this recording

  2. 9 years earlier
  3. AV

    Ashish Vaswani quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 1st of 24 in this piece

  4. NS

    Noam Shazeer quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 2nd of 24 in this piece

  5. NP

    Niki Parmar quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 3rd of 24 in this piece

  6. JU

    Jakob Uszkoreit quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 4th of 24 in this piece

  7. LJ

    Llion Jones quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 5th of 24 in this piece

  8. AG

    Aidan N. Gomez quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 6th of 24 in this piece

  9. ŁK

    Łukasz Kaiser quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 7th of 24 in this piece

  10. IP

    Illia Polosukhin quoted

    We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 8th of 24 in this piece

  11. AV

    Ashish Vaswani quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 9th of 24 in this piece

  12. NS

    Noam Shazeer quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 10th of 24 in this piece

  13. NP

    Niki Parmar quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 11th of 24 in this piece

  14. JU

    Jakob Uszkoreit quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 12th of 24 in this piece

  15. LJ

    Llion Jones quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 13th of 24 in this piece

  16. AG

    Aidan N. Gomez quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 14th of 24 in this piece

  17. ŁK

    Łukasz Kaiser quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 15th of 24 in this piece

  18. IP

    Illia Polosukhin quoted

    To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 16th of 24 in this piece

  19. AV

    Ashish Vaswani quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 17th of 24 in this piece

  20. NS

    Noam Shazeer quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 18th of 24 in this piece

  21. NP

    Niki Parmar quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 19th of 24 in this piece

  22. JU

    Jakob Uszkoreit quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 20th of 24 in this piece

  23. LJ

    Llion Jones quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 21st of 24 in this piece

  24. AG

    Aidan N. Gomez quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 22nd of 24 in this piece

  25. ŁK

    Łukasz Kaiser quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 23rd of 24 in this piece

  26. IP

    Illia Polosukhin quoted

    As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 24th of 24 in this piece