korrents

What Łukasz Kaiser thinks about transformers

@lukasz-kaiser · 3 positions · 0 changes of mind

One of the eight authors of the 2017 transformer paper "Attention Is All You Need", then at Google Brain; per the paper's own contribution note, designed and implemented parts of tensor2tensor, the codebase that replaced the team's earlier one.

Łukasz Kaiser did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

3 dated positions, 2017, in their own words. Our reading of what Łukasz Kaiser has said — not written or endorsed by them.

3 positions so far — this page is not yet offered to search engines.

  1. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.

    Attention is all you needarxiv.org 7th of 24 in this piece

  2. To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution.

    Attention is all you needarxiv.org 15th of 24 in this piece

  3. As side benefit, self-attention could yield more interpretable models.

    Attention is all you needarxiv.org 23rd of 24 in this piece