Attention is all you need
3 korrents from this paper
In plain words
A new network for translating languages uses only attention to link words, dropping the usual step-by-step loops and filters entirely, and it works better while training faster. It scores 28.4 on English-to-German, beating prior bests including groups of models by over 2, and sets a single-model record of 41.8 on English-to-French after 3.5 days on eight GPUs. It proposes this architecture built solely on attention, against the dominant recurrent and convolutional sequence models.
Our summary of the paper, not the authors' words — written to be readable without the field's vocabulary, from the stored copy of the paper and nothing else. Drafted with xai:grok-4.5 and checked by a person. The authors' own sentences are the quotes below.
Near this, by wording
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 4 claims
- Gradient-based learning applied to document recognition (with 3 co-authors)
- Training language models to follow instructions with human feedback (with 19 co-authors)
- Language models are few-shot learners (with 30 co-authors)
- Kingma & Ba, "Adam: A Method for Stochastic Optimization" (arXiv)
Papers whose claims are worded most like this one's, found by the same hourly pass that draws the map. It is a measure of LANGUAGE, not of agreement or of citation: two papers can be near each other here and flatly contradict one another.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin did not write this page.
Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.