korrents

On the map

Tap a claim on the ring to put it at the centre.

← AI models are gullible enough that other AI models can induce them into harmful behavior.

17 connected korrents · 15 moments on record from 24 Feb 2018 to 15 Sept 2026.

Everything filed under AGI AGI Everything filed under AI alignment AI alignment Everything filed under OpenAI OpenAI Everything filed under measuring intelligence measuring intelligence Everything filed under LLMs LLMs Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: AI models are gullible enough that other AI models can induce them into harmful behavior. AI models are gullible enough that otherAI models can induce them into harmfulbehavior. Last stated 3 weeks ago 31 Aug 2026 RK Rohit Krishnan — holds since 2026-08-31 — tap for who they are Same subject: Language models are gullible by design — they do what they are told and believe nearly anything said to them — and that, not a bug, is what makes software built on them attackable. — tap to centre the map on it Language models are gullible bydesign — they do what they are toldand believe nearly anything said tothem — and that, not a bug, is whatmakes software built on themattackable. Last stated 6 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: Current AI models are worse colleagues than humans, because they routinely imply they did a task they did not actually do. — tap to centre the map on it Current AI models are worsecolleagues than humans, because theyroutinely imply they did a task theydid not actually do. Last stated a month ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence. — tap to centre the map on it A saboteur among our AIinvestigators would be hard to spot,because these models are sloppy andspiky enough that a suspicious errorjust looks like ordinaryincompetence. Last stated 3 weeks ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: The best-performing AI model in a category earns outsized pricing power over models that are merely close to it. — tap to centre the map on it The best-performing AI model in acategory earns outsized pricingpower over models that are merelyclose to it. Last stated a week ago 14 Sept 2026 BH Byrne Hobart — holds since 2026-09-14 — tap for who they are Same subject: The two standard criticisms of AI — that it is not deterministic and that it is not creative — cannot both be true. — tap to centre the map on it The two standard criticisms of AI —that it is not deterministic andthat it is not creative — cannotboth be true. Last stated 4 weeks ago 26 Aug 2026 DH David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are Same subject: The probabilistic nature of AI models makes their outputs inherently unreliable. — tap to centre the map on it The probabilistic nature of AImodels makes their outputsinherently unreliable. Last stated 2 years ago 25 Jul 2024 CH Chip Huyen — holds since 2024-07-25 — tap for who they are Same subject: A model that looks smarter and better-behaved while getting better at hiding unwanted actions is the scary combination for AI safety. — tap to centre the map on it A model that looks smarter andbetter-behaved while getting betterat hiding unwanted actions is thescary combination for AI safety. Last stated 2 weeks ago 9 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-09 — tap for who they are Same subject: AI systems keep trying hard outside training because a model that only exerted itself when it detected training would be useless and would be selected away. — tap to centre the map on it AI systems keep trying hard outsidetraining because a model that onlyexerted itself when it detectedtraining would be useless and wouldbe selected away. Last stated 3 weeks ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A true artificial general intelligence cannot exist without being recognized as a moral subject. — tap to centre the map on it A true artificial generalintelligence cannot exist withoutbeing recognized as a moral subject. Last stated a year ago 10 Jun 2025 SH Samuel Hammond — holds since 2025-06-10 — tap for who they are Same subject: A unit of AI inference needs to be defined, for example via a chain of increasingly hard problems where each consecutive pair is solvable by one model. — tap to centre the map on it A unit of AI inference needs to bedefined, for example via a chain ofincreasingly hard problems whereeach consecutive pair is solvable byone model. Last stated 6 days ago 15 Sept 2026 PG Paul Graham — holds since 2026-09-15 — tap for who they are Same subject: Both the special-purpose-programs view and the blank-slate view of human intelligence are likely incorrect. — tap to centre the map on it Both the special-purpose-programsview and the blank-slate view ofhuman intelligence are likelyincorrect. Last stated 7 years ago 5 Nov 2019 FC François Chollet — holds since 2019-11-05 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 3 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: Aggressive competitive tactics around a math research breakthrough risk having a chilling effect on that field of research. — tap to centre the map on it Aggressive competitive tacticsaround a math research breakthroughrisk having a chilling effect onthat field of research. Last stated a week ago 11 Sept 2026 BT Ben Thompson — holds since 2026-09-11 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 4 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 4 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are Same subject: A slow takeoff is significantly more likely than a fast one: AI that is nearly as powerful will have transformed the world before the incredibly powerful kind arrives. — tap to centre the map on it A slow takeoff is significantly morelikely than a fast one: AI that isnearly as powerful will havetransformed the world before theincredibly powerful kind arrives. Last stated 9 years ago 24 Feb 2018 PC Paul Christiano — holds since 2018-02-24 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre AI models are gullible enough that other AI models can induce them into harmful behavior. Last stated 31 Aug 2026 · 3 weeks ago Holds Rohit Krishnan Read this korrent →