korrents

korrents · piece

Jan 2021 Gwern.net Newsletter

Gwern · 4 Feb 2021

The piece opens with

Mixture-of-experts models can scale stably but begin to show fundamental limits as they are scaled further.

the mixture-of-experts approach, while scaling stably, starts showing its limits

Read the piece on gwern.substack.com ↗

3 korrents from this piece ↓

Gwern did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. the mixture-of-experts approach, while scaling stably, starts showing its limits
  2. To take the GAN/DRL analogy seriously, perhaps they were they ultimately a dead end, akin to trying to learn everything from rewards, and an adversarial GAN loss ought to be only the cherry on the cake of a large unsupervised/semi-supervised generative model.
  3. The fact that the prefix-tuning, by directly optimizing the prompt embeddings, yields better performance than even single optimized text prompts, suggests so.

Keep reading

The other pieces are the ones whose claims are worded most like this one's. That measures language, not agreement: they may argue the opposite.