korrents
Sebastian Ruder

Sebastian Ruder

@sebastian-ruder · 23 positions · 0 changes of mind

Research scientist at Meta working on multilingual models and evaluation; led the multilingual team at Cohere and was at Google DeepMind before that.

Everything they publish, on ppll ↗

Sebastian Ruder did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

  1. LLMsbenchmarksscaling laws

    GPT models perform much better on coding problems released before their pre-training data cut-off.

    The Evolving Landscape of LLM Evaluationruder.io 1st of 2 in this piece

  2. benchmarks

    The time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.

    The Evolving Landscape of LLM Evaluationruder.io 2nd of 2 in this piece

  3. 4 weeks earlier
  4. LLMs

    I'm excited about what this means for the open-source community and research, with the gap between closed-source and open-weight models closing and SOTA-level conversational models being more easily accessible.

    Command R+ruder.io 1st of 2 in this piece

  5. LLMs

    Retrieval-augmented generation (RAG; Lewis et al., 2020), which conditions on the LLM's generation on retrieved documents is the most practical paradigm IMO.

    Command R+ruder.io 2nd of 2 in this piece

  6. 7 weeks earlier
  7. many under-represented languages are primarily spoken; they do not have a written tradition or a standardized orthography. Text-based NLP technology is thus of limited use to them.

    True Zero-shot MTruder.io 1st of 2 in this piece

  8. the embodied, interactive, and multi-modal nature of first language (L1) acquisition is challenging to replicate with current models.

    True Zero-shot MTruder.io 2nd of 2 in this piece

  9. 2 weeks earlier
  10. startups

    In light of the applied nature of current research problems, there is another path that exposes you to cutting-edge AI work: joining a startup.

    Thoughts on the 2024 AI Job Marketruder.io 1st of 3 in this piece

  11. This lack of knowledge sharing may impede progress in AI development.

    Thoughts on the 2024 AI Job Marketruder.io 2nd of 3 in this piece

  12. Such size poses challenges for the effective execution and prioritization, increasing friction and making it more difficult to quickly make decisions.

    Thoughts on the 2024 AI Job Marketruder.io 3rd of 3 in this piece

  13. 4 weeks earlier
  14. At its best, research is a collaborative and sometimes argumentative conversation.

    The Big Picture of AI Researchruder.io

  15. 4 weeks earlier
  16. while massive compute often achieves breakthrough results, its usage is often inefficient.

    NLP Research in the Era of LLMsruder.io 1st of 3 in this piece

  17. LLMs

    recent LLMs are reaching the limits of text data online and repeating data eventually leads to diminishing returns

    NLP Research in the Era of LLMsruder.io 2nd of 3 in this piece

  18. In the near term, the largest models using the most compute will continue to be the most capable.

    NLP Research in the Era of LLMsruder.io 3rd of 3 in this piece

  19. 2 weeks earlier
  20. LLMs

    LLMs are still limited in non-English settings and that making LLMs more multilingual is an important direction.

    EMNLP 2023 Primerruder.io 1st of 3 in this piece

  21. LLMs

    In light of the increasing number of closed-source LLMs, it is important to continue to promote an open culture of sharing knowledge, data, and software, from which the NLP community has benefited greatly.

    EMNLP 2023 Primerruder.io 2nd of 3 in this piece

  22. LLMs

    Full fine-tuning of LLMs has become prohibitive and requires parameter-efficient methods instead.

    EMNLP 2023 Primerruder.io 3rd of 3 in this piece

  23. 4 days earlier
  24. scaling laws

    In sum, whenever we don't have infinite amounts of pre-training data, we should train smaller models for more (up to 4) epochs.

    NeurIPS 2023 Primerruder.io 1st of 3 in this piece

  25. LLMs

    Overall, while initial plans of LLMs can be useful as a starting point, LLM-based planning currently works best mainly in conjunction with external tools.

    NeurIPS 2023 Primerruder.io 2nd of 3 in this piece

  26. LLMs

    Overall, current LLMs still struggle with composing operations into correct reasoning paths.

    NeurIPS 2023 Primerruder.io 3rd of 3 in this piece

  27. 2 weeks earlier
  28. training on a small set of high-quality data outperforms instruction-tuning on larger, noisier data.

    An Overview of Instruction Tuning Dataruder.io 1st of 2 in this piece

  29. OpenAI

    Models that are instruction-tuned on ChatGPT-generated data mimic ChatGPT's style (and may thus fool human raters!) but not its factuality

    An Overview of Instruction Tuning Dataruder.io 2nd of 2 in this piece

  30. 9 months earlier
  31. scaling laws

    Given the trend of pre-training larger and larger models, we believe modularity will be crucial.

    Modular Deep Learningruder.io 1st of 2 in this piece

  32. But modularity may also facilitate a shift away from a concentration of model development in a few institutions and to distributing the development of modular components across the community.

    Modular Deep Learningruder.io 2nd of 2 in this piece