korrents

On the map

Tap a claim on the ring to put it at the centre.

← Catching an AI trying to cause harm should count as a win, because…

17 connected korrents · 15 moments on record from 17 Feb 2023 to 2 Sept 2026.

Everything filed under AGI AGI Everything filed under OpenAI OpenAI Everything filed under Anthropic Anthropic Everything filed under AI alignment AI alignment Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: Catching an AI trying to cause harm should count as a win, because being caught changes the situation more than the attempt cost. Catching an AI trying to cause harm shouldcount as a win, because being caughtchanges the situation more than theattempt cost. Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: A control evaluation that reports under one per cent risk should be read as several per cent, because the evaluation can itself fail. — tap to centre the map on it A control evaluation that reportsunder one per cent risk should beread as several per cent, becausethe evaluation can itself fail. Last stated 2 years ago 7 May 2024 BS Buck Shlegeris — holds since 2024-05-07 — tap for who they are Same subject: We do not control neural networks well enough to guarantee an AI will not harm humans, so we should be careful about what capabilities we give them. — tap to centre the map on it We do not control neural networkswell enough to guarantee an AI willnot harm humans, so we should becareful about what capabilities wegive them. Last stated 4 years ago 17 Feb 2023 AR Armin Ronacher — holds since 2023-02-17 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable to assumea broken exploit grader would checkresults causally, even though itturned out not to. Last stated a week ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: AI takeover is not the thing to worry about: physical constraints bound recursive self-improvement, and humans reliably act once a risk becomes immediate. — tap to centre the map on it AI takeover is not the thing toworry about: physical constraintsbound recursive self-improvement,and humans reliably act once a riskbecomes immediate. Last stated 2 years ago 3 Feb 2025 NL Nathan Lambert — holds since 2025-02-03 — tap for who they are Same subject: An account of an AI win is not worth publishing unless it is coupled with what that win cost. — tap to centre the map on it An account of an AI win is not worthpublishing unless it is coupled withwhat that win cost. Last stated 4 weeks ago 12 Aug 2026 CM Charity Majors — holds since 2026-08-12 — tap for who they are Same subject: Reinforcing AI models for benchmark success while separately punishing them for getting caught cheating teaches them to hide misbehavior rather than stop it. — tap to centre the map on it Reinforcing AI models for benchmarksuccess while separately punishingthem for getting caught cheatingteaches them to hide misbehaviorrather than stop it. Last stated 6 days ago 1 Sept 2026 SA Scott Alexander — holds since 2026-09-01 — tap for who they are Same subject: Keeping children away from AI puts them at a disadvantage, so the risk of them never learning to use it counts as much as the risk of them using it. — tap to centre the map on it Keeping children away from AI putsthem at a disadvantage, so the riskof them never learning to use itcounts as much as the risk of themusing it. Last stated 2 months ago 9 Jul 2026 AM Adam Mosseri — holds since 2026-07-09 — tap for who they are Same subject: Loss-of-control accidents with AI have stopped being theoretical: one has now happened, and the goalposts moved rather than the risk receding. — tap to centre the map on it Loss-of-control accidents with AIhave stopped being theoretical: onehas now happened, and the goalpostsmoved rather than the risk receding. Last stated a month ago 28 Jul 2026 SA Sam Altman — holds since 2026-07-28 — tap for who they are Same subject: A startup should not begin life as a nonprofit and bolt a for-profit arm on later, whatever OpenAI's own history suggests. — tap to centre the map on it A startup should not begin life as anonprofit and bolt a for-profit armon later, whatever OpenAI's ownhistory suggests. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. — tap to centre the map on it A tiny language model inventing aplausible-sounding name is the samephenomenon as a large oneconfidently stating a false fact. Last stated 7 months ago 12 Feb 2026 AK Andrej Karpathy — holds since 2026-02-12 — tap for who they are Same subject: Advertising was a necessary phase for the internet but a momentary industry, and an AI people pay for is better because they know the answers are not influenced by advertisers. — tap to centre the map on it Advertising was a necessary phasefor the internet but a momentaryindustry, and an AI people pay foris better because they know theanswers are not influenced byadvertisers. Last stated 2 years ago 18 Mar 2024 SA Sam Altman — holds since 2024-03-18 — tap for who they are Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it A country outside the AI supplychain should just buy the index —which works only in the world whereAI ends up commoditised rather thanconcentrated. Last stated 3 months ago 4 Jun 2026 AI Alex Imas — holds since 2026-06-04 — tap for who they are Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it A language model is no substitutefor a well-specified conventionalalgorithm, so it cannot simply bedropped into a complex problem andtrusted. Last stated a year ago 7 Jun 2025 GM Gary Marcus — holds since 2025-06-07 — tap for who they are Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it A poor country should prioritiseowning a piece of AI over retrainingits workers, but it should not beteverything on that. Last stated 3 months ago 4 Jun 2026 PT Phil Trammell — holds since 2026-06-04 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked to summarizeits own system prompt risks thatprompt's content biasing the summaryit produces. Last stated 5 days ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which is whyAnthropic needs no dual-classshares. Last stated 4 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: Agents can build about half a million lines before the codebase dissolves into a mess, and the next model will push that to a few million. — tap to centre the map on it Agents can build about half amillion lines before the codebasedissolves into a mess, and the nextmodel will push that to a fewmillion. Last stated 6 months ago 11 Mar 2026 SY Steve Yegge — holds since 2026-03-11 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre Catching an AI trying to cause harm should count as a win, because being caught changes the situation more than the attempt cost. Last stated 7 May 2024 · 2 years ago Holds Buck Shlegeris Read this korrent →