korrents

korrents · piece

NeurIPS 2023 Primer

Sebastian Ruder · 1 Dec 2023 · ruder.io

3 korrents from this piece

Sebastian Ruder did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. In sum, whenever we don't have infinite amounts of pre-training data, we should train smaller models for more (up to 4) epochs.
  2. Overall, while initial plans of LLMs can be useful as a starting point, LLM-based planning currently works best mainly in conjunction with external tools.
  3. Overall, current LLMs still struggle with composing operations into correct reasoning paths.