korrents

A korrentour readingWhat is a korrent?

Current techniques restrict pre-trained representation power because standard language models are unidirectional.

Drawn from what Kristina Toutanova, Kenton Lee and 2 others said

What this subject means

scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.

What they actually said

Word for word, with the source under each one. They did not write this page.

  1. Kristina Toutanova We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training. arxiv.org

    One of the four authors of the 2018 paper "BERT:…

    We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
  2. Kenton Lee We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training. arxiv.org

    One of the four authors of the 2018 paper "BERT:…

    We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
  3. Ming-Wei Chang We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training. arxiv.org

    One of the four authors of the 2018 paper "BERT:…

    We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.
  4. Jacob Devlin We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training. arxiv.org

    One of the four authors of the 2018 paper "BERT:…

    We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.

Added to korrents 11 Oct 2018 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

On the map

Loading the map… or open it on its own page

Open the map on its own page →