Noam Shazeer
One of the eight authors of the 2017 transformer paper "Attention Is All You Need", then at Google Brain; per the paper's own contribution note, proposed scaled dot-product attention, multi-head attention and the parameter-free position representation.
Noam Shazeer did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
-
Their wordsWe propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
-
Their wordsTo the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.
↗Attention is all you needarxiv.org 10th of 24 in this piece
-
Their wordsAs side benefit, self-attention could yield more interpretable models.
↗Attention is all you needarxiv.org 18th of 24 in this piece