korrents

What Jan Leike thinks about LLMs

@jan-leike · 16 positions · 0 changes of mind

Alignment researcher. Co-led OpenAI's superalignment team until 2024 and has led alignment work at Anthropic since; writes Musings on the Alignment Problem.

Jan Leike did not write this page.

We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it. Or tell us what is wrong here.

4 dated positions, 2022, in their own words. Our reading of what Jan Leike has said — not written or endorsed by them.

4 positions so far — this page is not yet offered to search engines.

  1. Making language models bigger does not inherently make them better at following a user's intent.

    Training language models to follow instructions with human feedback (with 19 co-authors)arxiv.org 1st of 4 in this piece

  2. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users.

    Training language models to follow instructions with human feedback (with 19 co-authors)arxiv.org 2nd of 4 in this piece

  3. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.

    Training language models to follow instructions with human feedback (with 19 co-authors)arxiv.org 3rd of 4 in this piece

  4. This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.

    Training language models to follow instructions with human feedback (with 19 co-authors)arxiv.org 4th of 4 in this piece

    AI alignment