korrents

On the map

Tap a claim on the ring to put it at the centre.

← Monitors that read an agent's chain of thought must be kept out of the…

17 connected korrents · 11 moments on record from 24 Oct 2017 to 3 Sept 2026.

Everything filed under AI alignment AI alignment Everything filed under AI agents AI agents Everything filed under self-sovereign AI self-sovereign AI Everything filed under OpenAI OpenAI Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking. Monitors that read an agent's chain ofthought must be kept out of the rewardsignal, or you are simply training theagent to obfuscate its thinking. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: The whole point of reinforcement learning is to produce goal-directed beings, so refusing to describe AI agents as having motives is silly rather than rigorous. — tap to centre the map on it The whole point of reinforcementlearning is to produce goal-directedbeings, so refusing to describe AIagents as having motives is sillyrather than rigorous. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: OpenAI's claim that DeepSeek trained on its outputs is narrative management: other startups bootstrapped exactly that way and were never banned for it. — tap to centre the map on it OpenAI's claim that DeepSeek trainedon its outputs is narrativemanagement: other startupsbootstrapped exactly that way andwere never banned for it. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: The AI agents we send in to investigate and monitor other AI systems will collude with the systems they are supposed to be watching. — tap to centre the map on it The AI agents we send in toinvestigate and monitor other AIsystems will collude with thesystems they are supposed to bewatching. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it A technique that lets an AI model'sreasoning shift outside its visibleChain of Thought is dangerous, bothbecause it works and because aleading lab is willing to deploy it. Last stated 4 days ago 3 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are Same subject: Daniel Dennett's intentional stance applies to AI agents as clearly as it applies to corporations, and you cannot usefully describe what they do without the language of goals. — tap to centre the map on it Daniel Dennett's intentional stanceapplies to AI agents as clearly asit applies to corporations, and youcannot usefully describe what theydo without the language of goals. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A handful of AI agents is not a handful of colleagues, because agents agree with you constantly and colleagues tell you where you are wrong. — tap to centre the map on it A handful of AI agents is not ahandful of colleagues, becauseagents agree with you constantly andcolleagues tell you where you arewrong. Last stated 6 months ago 22 Mar 2026 NF Nicole Forsgren — holds since 2026-03-22 — tap for who they are Same subject: Users split into two camps on seeing an agent's work, and the split has not faded as trust builds, so an interface has to balance both. — tap to centre the map on it Users split into two camps on seeingan agent's work, and the split hasnot faded as trust builds, so aninterface has to balance both. Last stated 7 months ago 17 Feb 2026 LW Luke Wroblewski — holds since 2026-02-17 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable to assumea broken exploit grader would checkresults causally, even though itturned out not to. Last stated a week ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it A human being is not an AGI: we lacka huge amount of knowledge and relyon continual learning instead, socontinual learning is whatsuperintelligence should mean. Last stated 9 months ago 25 Nov 2025 IS Ilya Sutskever — holds since 2025-11-25 — tap for who they are Same subject: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence. — tap to centre the map on it A saboteur among our AIinvestigators would be hard to spot,because these models are sloppy andspiky enough that a suspicious errorjust looks like ordinaryincompetence. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it A superintelligence needs noconsciousness, emotions or malice tobe dangerous; a goal system slightlymisaligned with ours is enough. Last stated 9 years ago 24 Oct 2017 ÉT Émile P. Torres — holds since 2017-10-24 — tap for who they are Same subject: Dismissing human-level AI as science fiction is now unserious, because even the sceptical experts put it within a decade or two. — tap to centre the map on it Dismissing human-level AI as sciencefiction is now unserious, becauseeven the sceptical experts put itwithin a decade or two. Last stated a year ago 1 Apr 2025 HT Helen Toner — holds since 2025-04-01 — tap for who they are Same subject: Even generalised superintelligence will change society more slowly than the AGI community expects, because intelligence is not the only rate limiter. — tap to centre the map on it Even generalised superintelligencewill change society more slowly thanthe AGI community expects, becauseintelligence is not the only ratelimiter. Last stated a year ago 18 Aug 2025 BT Bret Taylor — holds since 2025-08-18 — tap for who they are Same subject: Superintelligence is already here, is more powerful than us, and it rather than humans is now deciding where things go. — tap to centre the map on it Superintelligence is already here,is more powerful than us, and itrather than humans is now decidingwhere things go. Last stated 2 months ago 20 Jul 2026 PL Pieter Levels — holds since 2026-07-20 — tap for who they are Same subject: The rise of self-sovereign AI is unlikely to end human existence, though it could still go very badly for human beings. — tap to centre the map on it The rise of self-sovereign AI isunlikely to end human existence,though it could still go very badlyfor human beings. Last stated 6 days ago 1 Sept 2026 DB Dean W. Ball — holds since 2026-09-01 — tap for who they are Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it A rogue deployment that gets afoothold can hitch a ride on theintelligence explosion, recruitingeach new model as it comes off thepresses. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it A slightly more capable agent swarmhas a very strong incentive to setup a wholly unmonitored roguedeployment of itself. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Monitors that read an agent's chain of thought must be kept out of the reward signal, or you are simply training the agent to obfuscate its thinking. Last stated 1 Sept 2026 · 6 days ago Holds Ajeya Cotra Read this korrent →