Tap a claim on the ring to put it at the centre.
← Deleting the rollouts where a monitor caught cheating is structurally…
12 connected korrents · 8 moments on record from 24 Oct 2017 to 3 Sept 2026.
Everything filed under AI alignment
AI alignment
Everything filed under benchmarks
benchmarks
Everything filed under AI agents
AI agents
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Deleting the rollouts where a monitor caught cheating is structurally the same as positively reinforcing all the cheating the monitor missed.
Deleting the rollouts where a monitor caught cheating is structurally the same as positively reinforcing all the cheating the monitor missed.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it. — tap to centre the map on it
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating…
Last stated 6 days ago
1 Sept 2026
SA
Scott Alexander — holds since 2026-09-01 — tap for who they are
Same subject: If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see. — tap to centre the map on it
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches…
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: No person learns the way RL does: a human reviews which parts of an attempt were good instead of rewarding every step of a lucky one. — tap to centre the map on it
No person learns the way RL does: a human reviews which parts of an attempt were good instead of rewarding every…
Last stated 11 months ago
17 Oct 2025
AK
Andrej Karpathy — holds since 2025-10-17 — tap for who they are
Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so…
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned…
Last stated 9 years ago
24 Oct 2017
ÉT
Émile P. Torres — holds since 2017-10-24 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both…
Last stated 4 days ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number. — tap to centre the map on it
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a…
Last stated 2 months ago
26 Jun 2026
NB
Noam Brown — holds since 2026-06-26 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a…
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost. — tap to centre the map on it
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own…
Last stated 5 months ago
1 Apr 2026
TP
Thuan Pham — holds since 2026-04-01 — tap for who they are
Same subject: Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth. — tap to centre the map on it
Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on…
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
Same subject: Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality. — tap to centre the map on it
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
Same subject: Many self-sovereign AI agents will fund themselves by committing or facilitating crime, because crime is high-margin work. — tap to centre the map on it
Many self-sovereign AI agents will fund themselves by committing or facilitating crime, because crime is…
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: how recently it was last stated — full and dark this week, a faint sliver at five years a face: someone on record holding the claim — tap it for who they are
At the centre
Deleting the rollouts where a monitor caught cheating is structurally the same as positively reinforcing all the cheating the monitor missed.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Read this korrent →
Similar wording
Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it.
Last stated 1 Sept 2026 · 6 days ago
Holds SA Scott Alexander
Similar wording
If whole categories of reward hack are undetectable by humans, then training against the hacks we do catch teaches models to cheat only where we cannot see.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
No person learns the way RL does: a human reviews which parts of an attempt were good instead of rewarding every step of a lucky one.
Last stated 17 Oct 2025 · 11 months ago
Holds Andrej Karpathy
Same subject: AI alignment
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Same subject: AI alignment
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 24 Oct 2017 · 9 years ago
Holds ÉT Émile P. Torres
Same subject: AI alignment
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 days ago
Holds ZM Zvi Mowshowitz
Same subject: benchmarks
A benchmark result should be reported under a stated budget, or as a curve against test-time compute — never as a single number.
Last stated 26 Jun 2026 · 2 months ago
Holds NB Noam Brown
Same subject: benchmarks
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: benchmarks
A company's staff-engineer bar should be set against the best companies in the industry rather than against its own history, which is what makes title inflation a real cost.
Last stated 1 Apr 2026 · 5 months ago
Holds TP Thuan Pham
Same subject: self-sovereign AI
Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball
Same subject: self-sovereign AI
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball
Same subject: self-sovereign AI
Many self-sovereign AI agents will fund themselves by committing or facilitating crime, because crime is high-margin work.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball