korrents

On the map

Tap a claim on the ring to put it at the centre.

← Very capable AI will be harder to align than current systems, because…

17 connected korrents · 10 moments on record from 10 Jun 2022 to 2 Sept 2026.

Everything filed under AI alignment AI alignment Everything filed under AGI AGI Everything filed under OpenAI OpenAI Everything filed under self-sovereign AI self-sovereign AI Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Very capable AI will be harder to align than current systems, because the loop of spotting a bad behaviour and patching the training that caused it breaks down. Very capable AI will be harder to alignthan current systems, because the loop ofspotting a bad behaviour and patching thetraining that caused it breaks down. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: Misaligned AI behaviour will keep getting rarer and, at the same time, keep getting more extreme. — tap to centre the map on it Misaligned AI behaviour will keepgetting rarer and, at the same time,keep getting more extreme. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: Current AI models are worse colleagues than humans, because they routinely imply they did a task they did not actually do. — tap to centre the map on it Current AI models are worsecolleagues than humans, because theyroutinely imply they did a task theydid not actually do. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: AI systems keep trying hard outside training because a model that only exerted itself when it detected training would be useless and would be selected away. — tap to centre the map on it AI systems keep trying hard outsidetraining because a model that onlyexerted itself when it detectedtraining would be useless and wouldbe selected away. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Building AI that can actually be trusted is the goal; containing an AI known to be misaligned is not a substitute for it. — tap to centre the map on it Building AI that can actually betrusted is the goal; containing anAI known to be misaligned is not asubstitute for it. Last stated 2 years ago 24 Jan 2025 JL Jan Leike — holds since 2025-01-24 — tap for who they are Same subject: The likeliest bad path is not a coup but sloppiness: AI does everything verifiable well, research races ahead, and the subtle work of keeping AI safe is what gets done badly. — tap to centre the map on it The likeliest bad path is not a coupbut sloppiness: AI does everythingverifiable well, research racesahead, and the subtle work ofkeeping AI safe is what gets donebadly. Last stated 4 weeks ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous. — tap to centre the map on it Relying on gradual, continuousshifts in AI training behavior tocatch misalignment will eventuallyfail because the dangerous shiftitself may be discontinuous. Last stated 5 days ago 2 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-02 — tap for who they are Same subject: Adapting over time is a workable answer to humans misusing AI and no answer at all to losing control of systems smarter than us. — tap to centre the map on it Adapting over time is a workableanswer to humans misusing AI and noanswer at all to losing control ofsystems smarter than us. Last stated a year ago 5 Apr 2025 HT Helen Toner — holds since 2025-04-05 — tap for who they are Same subject: Aligning an AI is not impossible in principle; the problem is that it has to be done without the textbook that would make it look easy. — tap to centre the map on it Aligning an AI is not impossible inprinciple; the problem is that ithas to be done without the textbookthat would make it look easy. Last stated 4 years ago 10 Jun 2022 EY Eliezer Yudkowsky — holds since 2022-06-10 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it A tiny language model inventing aplausible-sounding name is the samephenomenon as a large oneconfidently stating a false fact. Last stated 7 months ago 12 Feb 2026 AK Andrej Karpathy — holds since 2026-02-12 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 3 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 3 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it A rogue deployment that gets afoothold can hitch a ride on theintelligence explosion, recruitingeach new model as it comes off thepresses. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it A slightly more capable agent swarmhas a very strong incentive to setup a wholly unmonitored roguedeployment of itself. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality. — tap to centre the map on it Banning all self-sovereign AI agentswould backfire, denying themlegitimate work and pushing theminto criminality. Last stated 6 days ago 1 Sept 2026 DB Dean W. Ball — holds since 2026-09-01 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Very capable AI will be harder to align than current systems, because the loop of spotting a bad behaviour and patching the training that caused it breaks down. Last stated 11 Aug 2026 · 4 weeks ago Holds Ryan Greenblatt Read this korrent →