korrents

A korrentour readingWhat is a korrent?

Humans keep their place in AI as judges rather than authors, because telling which of two answers is better is far easier than writing a good one.

Drawn from what Nathan Lambert said

What this subject means

reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.

A private bookmark. Not a position, and never counted.

What Nathan Lambert actually said

Word for word, with the source under each one. They did not write this page.

  1. Nathan Lambert

    Research scientist at the Allen Institute for AI

    And humans are actually very good at reading or judging between two things versus... This goes back to the core of what RLHF and preference tuning is that it's hard to generate a good answer for a lot of problems, but it's easy to see which one is better.

Added to korrents 3 Feb 2025 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

On the map

Loading the map… or open it on its own page

Open the map on its own page →