Tap a claim on the ring to put it at the centre.
← Reinforcement learning from verifiable rewards does not reliably carry…
17 connected korrents · 17 moments on record from 1 Oct 2023 to 13 Sept 2026. Nearly all of them are about reinforcement learning .
Everything filed under scaling laws
scaling laws
Everything filed under robotics
robotics
Everything filed under benchmarks
benchmarks
Everything filed under world models
world models
Everything filed under LLMs
LLMs
Everything filed under neural networks
neural networks
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Last stated a month ago
11 Aug 2026
SH
Samuel Hammond — holds since 2026-08-11 — tap for who they are
Same subject: Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training. — tap to centre the map on it
Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training.
Last stated 10 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for. — tap to centre the map on it
One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for.
Last stated a month ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware. — tap to centre the map on it
Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: RL environments exist to make a model generalise, not to teach it each skill one at a time — exactly as pre-training does. — tap to centre the map on it
RL environments exist to make a model generalise, not to teach it each skill one at a time — exactly as pre-training does.
Last stated 7 months ago
13 Feb 2026
DA
Dario Amodei — holds since 2026-02-13 — tap for who they are
Same subject: A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well. — tap to centre the map on it
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 10 months ago
17 Nov 2025
AK
Andrej Karpathy — holds since 2025-11-17 — tap for who they are
Same subject: Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago. — tap to centre the map on it
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: Current models are unsafe because they are not smart enough, not because they are too smart. — tap to centre the map on it
Current models are unsafe because they are not smart enough, not because they are too smart.
Last stated a week ago
13 Sept 2026
FC
François Chollet — holds since 2026-09-13 — tap for who they are
Same subject: Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving. — tap to centre the map on it
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 weeks ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old. — tap to centre the map on it
A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old.
Last stated 3 years ago
19 Dec 2023
PK
Philip Koopman — holds since 2023-12-19 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 6 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: A model that can predict the next frames of a video coherently understands the world, in the only sense of the word that is doing any work. — tap to centre the map on it
A model that can predict the next frames of a video coherently understands the world, in the only sense of the word that is doing any work.
Last stated a year ago
23 Jul 2025
DH
Demis Hassabis — holds since 2025-07-23 — tap for who they are
Same subject: An adaptive agent's world model cannot be a static representation learned once and fixed — tap to centre the map on it
An adaptive agent's world model cannot be a static representation learned once and fixed
Last stated a year ago
17 Jul 2025
FC
François Chollet — holds since 2025-07-17 — tap for who they are
Same subject: Coding is not just another application domain for AI; it is the meta-skill that lets AI generate its own training material and start the recursive self-improvement loop. — tap to centre the map on it
Coding is not just another application domain for AI; it is the meta-skill that lets AI generate its own training material and start the recursive self-improvement loop.
Last stated a month ago
10 Aug 2026
FC
François Chollet — holds since 2026-08-10 — tap for who they are
Same subject: A highly dynamic economy that cheaply mass-produces robots may increase rather than decrease the risk of Malthusian resource dilemmas. — tap to centre the map on it
A highly dynamic economy that cheaply mass-produces robots may increase rather than decrease the risk of Malthusian resource dilemmas.
Last stated 3 years ago
1 Oct 2023
TC
Tyler Cowen — holds since 2023-10-01 — tap for who they are
Same subject: A humanoid robot is the wrong tool on a factory floor, because a machine specialised in the task beats it. — tap to centre the map on it
A humanoid robot is the wrong tool on a factory floor, because a machine specialised in the task beats it.
Last stated a year ago
6 Apr 2025
MH
Molson Hart — holds since 2025-04-06 — tap for who they are
Same subject: A physical AI agent has to be safe on day one, because you cannot ship something merely good enough and let users find the edge cases. — tap to centre the map on it
A physical AI agent has to be safe on day one, because you cannot ship something merely good enough and let users find the edge cases.
Last stated 2 months ago
3 Aug 2026
DD
Dmitri Dolgov — holds since 2026-08-03 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Reinforcement learning from verifiable rewards does not reliably carry a model's pretrained ethical understanding into its actual behavior.
Last stated 11 Aug 2026 · a month ago
Holds Samuel Hammond
Read this korrent →
Same subject: scaling laws, reinforcement learning
Models resemble each other because pre-training is the same everywhere; what differentiates labs now is RL and post-training.
Last stated 25 Nov 2025 · 10 months ago
Holds Ilya Sutskever
Same subject: scaling laws, reinforcement learning
One pre-trained robot model now matches or beats the specialists that were fine-tuned with reinforcement learning for the very tasks they were built for.
Last stated 12 Aug 2026 · a month ago
Holds Chelsea Finn
Same subject: scaling laws, reinforcement learning
Pre-training is a crappy evolution: the practically buildable substitute for the process that gave animals their built-in hardware.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: scaling laws, reinforcement learning
RL environments exist to make a model generalise, not to teach it each skill one at a time — exactly as pre-training does.
Last stated 13 Feb 2026 · 7 months ago
Holds Dario Amodei
Same subject: reinforcement learning
A verifiable task can be optimised by reinforcement learning until a neural network performs it extremely well.
Last stated 17 Nov 2025 · 10 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Capability does not generalise for free: a model that will move mountains on an agentic task still tells the same bad joke it told five years ago.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: reinforcement learning
Current models are unsafe because they are not smart enough, not because they are too smart.
Last stated 13 Sept 2026 · a week ago
Holds François Chollet
Same subject: reinforcement learning
Humans barely use reinforcement learning for intelligence — what RL they do use goes into motor tasks, not problem solving.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 3 weeks ago
Holds Dax Raad
Same subject: benchmarks
A carmaker's claim to be the safest is mostly an artefact of comparing a new car against a fleet average twelve years old.
Last stated 19 Dec 2023 · 3 years ago
Holds Philip Koopman
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 6 months ago
Holds Thuan Pham
Same subject: world models
A model that can predict the next frames of a video coherently understands the world, in the only sense of the word that is doing any work.
Last stated 23 Jul 2025 · a year ago
Holds Demis Hassabis
Same subject: world models
An adaptive agent's world model cannot be a static representation learned once and fixed
Last stated 17 Jul 2025 · a year ago
Holds François Chollet
Same subject: world models
Coding is not just another application domain for AI; it is the meta-skill that lets AI generate its own training material and start the recursive self-improvement loop.
Last stated 10 Aug 2026 · a month ago
Holds François Chollet
Same subject: robotics
A highly dynamic economy that cheaply mass-produces robots may increase rather than decrease the risk of Malthusian resource dilemmas.
Last stated 1 Oct 2023 · 3 years ago
Holds Tyler Cowen
Same subject: robotics
A humanoid robot is the wrong tool on a factory floor, because a machine specialised in the task beats it.
Last stated 6 Apr 2025 · a year ago
Holds Molson Hart
Same subject: robotics
A physical AI agent has to be safe on day one, because you cannot ship something merely good enough and let users find the edge cases.
Last stated 3 Aug 2026 · 2 months ago
Holds Dmitri Dolgov