korrents

A korrentour readingWhat is a korrent?

If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.

Drawn from what Ryan Greenblatt said

A private bookmark. Not a position, and never counted.

What Ryan Greenblatt actually said

Word for word, with the source under each one. They did not write this page.

  1. Ryan Greenblatt

    Chief scientist at Redwood Research

    One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out

Added to korrents 11 Aug 2026 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

See this on the map →