Sebastian Ruder
Research scientist at Meta working on multilingual models and evaluation; led the multilingual team at Cohere and was at Google DeepMind before that.
Everything they publish, on ppll ↗
Sebastian Ruder did not write this page.
We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.
-
Their wordsGPT models perform much better on coding problems released before their pre-training data cut-off.
↗The Evolving Landscape of LLM Evaluationruder.io 1st of 2 in this piece
-
Their wordsThe time when benchmarks lasted multiple decades has passed. Going forward, we will rely less on public benchmark results.
↗The Evolving Landscape of LLM Evaluationruder.io 2nd of 2 in this piece
- 4 weeks earlier
-
Their wordsI'm excited about what this means for the open-source community and research, with the gap between closed-source and open-weight models closing and SOTA-level conversational models being more easily accessible.
-
Our reading
Retrieval-augmented generation is the most practical approach for reducing LLM hallucinations.
Their wordsRetrieval-augmented generation (RAG; Lewis et al., 2020), which conditions on the LLM's generation on retrieved documents is the most practical paradigm IMO.
- 7 weeks earlier
-
Their wordsmany under-represented languages are primarily spoken; they do not have a written tradition or a standardized orthography. Text-based NLP technology is thus of limited use to them.
-
Their wordsthe embodied, interactive, and multi-modal nature of first language (L1) acquisition is challenging to replicate with current models.
- 2 weeks earlier
-
Their wordsIn light of the applied nature of current research problems, there is another path that exposes you to cutting-edge AI work: joining a startup.
↗Thoughts on the 2024 AI Job Marketruder.io 1st of 3 in this piece
-
Their wordsThis lack of knowledge sharing may impede progress in AI development.
↗Thoughts on the 2024 AI Job Marketruder.io 2nd of 3 in this piece
-
Their wordsSuch size poses challenges for the effective execution and prioritization, increasing friction and making it more difficult to quickly make decisions.
↗Thoughts on the 2024 AI Job Marketruder.io 3rd of 3 in this piece
- 4 weeks earlier
-
Their wordsAt its best, research is a collaborative and sometimes argumentative conversation.
- 4 weeks earlier
-
Our reading
Breakthroughs achieved with massive compute are typically achieved using it inefficiently.
Their wordswhile massive compute often achieves breakthrough results, its usage is often inefficient.
↗NLP Research in the Era of LLMsruder.io 1st of 3 in this piece
-
Their wordsrecent LLMs are reaching the limits of text data online and repeating data eventually leads to diminishing returns
↗NLP Research in the Era of LLMsruder.io 2nd of 3 in this piece
-
Their wordsIn the near term, the largest models using the most compute will continue to be the most capable.
↗NLP Research in the Era of LLMsruder.io 3rd of 3 in this piece
- 2 weeks earlier
-
Their wordsLLMs are still limited in non-English settings and that making LLMs more multilingual is an important direction.
-
Their wordsIn light of the increasing number of closed-source LLMs, it is important to continue to promote an open culture of sharing knowledge, data, and software, from which the NLP community has benefited greatly.
-
Their wordsFull fine-tuning of LLMs has become prohibitive and requires parameter-efficient methods instead.
- 4 days earlier
-
Their wordsIn sum, whenever we don't have infinite amounts of pre-training data, we should train smaller models for more (up to 4) epochs.
-
Their wordsOverall, while initial plans of LLMs can be useful as a starting point, LLM-based planning currently works best mainly in conjunction with external tools.
-
Their wordsOverall, current LLMs still struggle with composing operations into correct reasoning paths.
- 2 weeks earlier
-
Their wordstraining on a small set of high-quality data outperforms instruction-tuning on larger, noisier data.
↗An Overview of Instruction Tuning Dataruder.io 1st of 2 in this piece
-
Their wordsModels that are instruction-tuned on ChatGPT-generated data mimic ChatGPT's style (and may thus fool human raters!) but not its factuality
↗An Overview of Instruction Tuning Dataruder.io 2nd of 2 in this piece
- 9 months earlier
-
Their wordsGiven the trend of pre-training larger and larger models, we believe modularity will be crucial.
-
Their wordsBut modularity may also facilitate a shift away from a concentration of model development in a few institutions and to distributing the development of modular components across the community.