korrents

A korrentour readingWhat is a korrent?

Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.

Drawn from what Samuel Hammond said

What these subjects mean

scaling laws Whether more compute keeps buying more capability, and where the curve now bends -- pre-training, post-training, test-time.

reinforcement learning Training by reward: environments, verifiable tasks, value functions, and whether it works or merely beats what came before.

What Samuel Hammond actually said

Word for word, with the source under each one. They did not write this page.

  1. Samuel Hammond

    Senior economist at the Foundation for American Innovation

    In short, RLVR gives us no reason to expect the normative representations learned in pre-training will acquire motivational force over a given action, especially when post-training repeatedly selects trajectories for terminal task success.

Added to korrents 11 Aug 2026 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

On the map

Loading the map… or open it on its own page

Open the map on its own page →