korrents
Sebastian Ruder

What Sebastian Ruder thinks about scaling laws

@sebastian-ruder · 23 positions · 0 changes of mind

Research scientist at Meta working on multilingual models and evaluation; led the multilingual team at Cohere and was at Google DeepMind before that.

Everything they publish, on ppll ↗

Sebastian Ruder did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

3 dated positions, 2023 to 2024, in their own words. Our reading of what Sebastian Ruder has said — not written or endorsed by them.

3 positions so far — this page is not yet offered to search engines.

  1. LLMsbenchmarks

    GPT models perform much better on coding problems released before their pre-training data cut-off.

    The Evolving Landscape of LLM Evaluationruder.io 1st of 2 in this piece

  2. 5 months earlier
  3. In sum, whenever we don't have infinite amounts of pre-training data, we should train smaller models for more (up to 4) epochs.

    NeurIPS 2023 Primerruder.io 1st of 3 in this piece

  4. 9 months earlier
  5. Given the trend of pre-training larger and larger models, we believe modularity will be crucial.

    Modular Deep Learningruder.io 1st of 2 in this piece