A korrentour readingWhat is a korrent?
A transduction model can rely entirely on self-attention for input and output representations without sequence-aligned RNNs or convolution.
Drawn from what Illia Polosukhin, Łukasz Kaiser and 6 others said
What this subject means
transformers The architecture that dropped recurrence in favour of attention, and that almost every large model since is built on.