korrents

korrents · piece

Why We Think

Lilian Weng · 1 May 2025 · lilianweng.github.io

3 korrents from this piece

Lilian Weng did not write this page.

Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.

  1. this self-correction capability turns out to not exist intrinsically among LLMs and does not easily work out of the box, due to various failure modes
  2. thinking for longer should be especially useful when the model is presented with an unusual input, such as an adversarial example or jailbreak attempt
  3. model CoTs could be biased due to lack of explicit training objectives aimed at encouraging faithful reasoning