korrents

Jacob Devlin

@jacob-devlin · 4 positions · 0 changes of mind

One of the four authors of the 2018 paper "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding", at Google AI Language, the affiliation printed on the paper.

Jacob Devlin did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

  1. We argue that current techniques restrict the power of the pre-trained representations, especially for the fine-tuning approaches. The major limitation is that standard language models are unidirectional, and this limits the choice of architectures that can be used during pre-training.

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understandingarxiv.org 1st of 16 in this piece

    LLMsscaling laws

  2. Such restrictions are sub-optimal for sentence-level tasks, and could be very harmful when applying fine-tuning based approaches to token-level tasks such as question answering, where it is crucial to incorporate context from both directions.

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understandingarxiv.org 5th of 16 in this piece

  3. We show that pre-trained representations reduce the need for many heavily-engineered task-specific architectures.

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understandingarxiv.org 9th of 16 in this piece

  4. Intuitively, it is reasonable to believe that a deep bidirectional model is strictly more powerful than either a left-to-right model or the shallow concatenation of a left-to-right and a right-to-left model.

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understandingarxiv.org 13th of 16 in this piece