Tap a claim on the ring to put it at the centre.
← Monitors that read an agent's chain of thought must be kept out of the…
17 connected korrents · 11 moments on record from 24 Oct 2017 to 3 Sept 2026.
Everything filed under AI alignment
AI alignment
Everything filed under AI agents
AI agents
Everything filed under self-sovereign AI
self-sovereign AI
Everything filed under OpenAI
OpenAI
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking.
Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous. — tap to centre the map on it
The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: OpenAI's claim that DeepSeek trained on its outputs is narrative management: other startups bootstrapped exactly that way and were never banned for it. — tap to centre the map on it
OpenAI's claim that DeepSeek trained on its outputs is narrative management: other startups bootstrapped exactly that way and were never banned for it.
Last stated 2 years ago
3 Feb 2025
NL
Nathan Lambert — holds since 2025-02-03 — tap for who they are
Same subject: The AI agents we send in to investigate and monitor other AI systems will collude with the systems they are supposed to be watching. — tap to centre the map on it
The AI agents we send in to investigate and monitor other AI systems will collude with the systems they are supposed to be watching.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 4 days ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: Daniel Dennett's intentional stance applies to AI agents as clearly as it applies to corporations, and you cannot usefully describe what they do without the language of goals. — tap to centre the map on it
Daniel Dennett's intentional stance applies to AI agents as clearly as it applies to corporations, and you cannot usefully describe what they do without the language of goals.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A handful of AI agents is not a handful of colleagues, because agents agree with you constantly and colleagues tell you where you are wrong. — tap to centre the map on it
A handful of AI agents is not a handful of colleagues, because agents agree with you constantly and colleagues tell you where you are wrong.
Last stated 6 months ago
22 Mar 2026
NF
Nicole Forsgren — holds since 2026-03-22 — tap for who they are
Same subject: Users split into two camps on seeing an agent's work, and the split has not faded as trust builds, so an interface has to balance both. — tap to centre the map on it
Users split into two camps on seeing an agent's work, and the split has not faded as trust builds, so an interface has to balance both.
Last stated 7 months ago
17 Feb 2026
LW
Luke Wroblewski — holds since 2026-02-17 — tap for who they are
Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated a week ago
29 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are
Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence. — tap to centre the map on it
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 9 years ago
24 Oct 2017
ÉT
Émile P. Torres — holds since 2017-10-24 — tap for who they are
Same subject: Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two. — tap to centre the map on it
Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two.
Last stated a year ago
1 Apr 2025
HT
Helen Toner — holds since 2025-04-01 — tap for who they are
Same subject: Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter. — tap to centre the map on it
Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter.
Last stated a year ago
18 Aug 2025
BT
Bret Taylor — holds since 2025-08-18 — tap for who they are
Same subject: Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go. — tap to centre the map on it
Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go.
Last stated 2 months ago
20 Jul 2026
PL
Pieter Levels — holds since 2026-07-20 — tap for who they are
Same subject: The rise of self-sovereign AI is unlikely to end human existence, though it could still go very badly for human beings. — tap to centre the map on it
The rise of self-sovereign AI is unlikely to end human existence, though it could still go very badly for human beings.
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Read this korrent →
Similar wording
The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
OpenAI's claim that DeepSeek trained on its outputs is narrative management: other startups bootstrapped exactly that way and were never banned for it.
Last stated 3 Feb 2025 · 2 years ago
Holds NL Nathan Lambert
Similar wording
The AI agents we send in to investigate and monitor other AI systems will collude with the systems they are supposed to be watching.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 days ago
Holds ZM Zvi Mowshowitz
Similar wording
Daniel Dennett's intentional stance applies to AI agents as clearly as it applies to corporations, and you cannot usefully describe what they do without the language of goals.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
A handful of AI agents is not a handful of colleagues, because agents agree with you constantly and colleagues tell you where you are wrong.
Last stated 22 Mar 2026 · 6 months ago
Holds NF Nicole Forsgren
Similar wording
Users split into two camps on seeing an agent's work, and the split has not faded as trust builds, so an interface has to balance both.
Last stated 17 Feb 2026 · 7 months ago
Holds LW Luke Wroblewski
Similar wording
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated 29 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Same subject: AI alignment
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Same subject: AI alignment
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: AI alignment
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 24 Oct 2017 · 9 years ago
Holds ÉT Émile P. Torres
Same subject: AGI
Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two.
Last stated 1 Apr 2025 · a year ago
Holds HT Helen Toner
Same subject: AGI
Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter.
Last stated 18 Aug 2025 · a year ago
Holds BT Bret Taylor
Same subject: AGI
Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go.
Last stated 20 Jul 2026 · 2 months ago
Holds Pieter Levels
Same subject: self-sovereign AI
The rise of self-sovereign AI is unlikely to end human existence, though it could still go very badly for human beings.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball
Same subject: self-sovereign AI
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra