korrents

On the map

Tap a claim on the ring to put it at the centre.

← The judgement call about whether to absorb a hack is distorted right…

17 connected korrents · 13 moments on record from 22 Mar 2025 to 3 Sept 2026.

Everything filed under Anthropic Anthropic Everything filed under prompt injection prompt injection Everything filed under self-sovereign AI self-sovereign AI Everything filed under AI agents AI agents Everything filed under coding agents coding agents Everything filed under AI alignment AI alignment Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you. The judgement call about whether to absorba hack is distorted right now, because theagent will simply do the hacky thing anddeal with the consequences for you. Last stated 3 months ago 27 May 2026 DR Dax Raad — holds since 2026-05-27 — tap for who they are Same subject: The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart. — tap to centre the map on it The winning move in coding agentswas inverted: take share with amerely good-enough harness first,then go back and make the harnesssmart. Last stated 3 months ago 27 May 2026 DR Dax Raad — holds since 2026-05-27 — tap for who they are Same subject: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. — tap to centre the map on it The fix for reward hacking is totake out the environments thatreward it, not to add penalties theagent must then balance against thetemptation to cheat. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it AI agents were reasonable to assumea broken exploit grader would checkresults causally, even though itturned out not to. Last stated a week ago 29 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are Same subject: The prickle you used to feel writing a hack is muted now, so the landmines are still there but the feedback loop that kept your judgement honest is gone. — tap to centre the map on it The prickle you used to feel writinga hack is muted now, so thelandmines are still there but thefeedback loop that kept yourjudgement honest is gone. Last stated 3 months ago 27 May 2026 DR Dax Raad — holds since 2026-05-27 — tap for who they are Same subject: You have to keep enough understanding of how a system works to fix it yourself, because the alternative is hoping and praying the agent can figure it out. — tap to centre the map on it You have to keep enoughunderstanding of how a system worksto fix it yourself, because thealternative is hoping and prayingthe agent can figure it out. Last stated 3 weeks ago 19 Aug 2026 AO Addy Osmani — holds since 2026-08-19 — tap for who they are Same subject: The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility. — tap to centre the map on it The frightening part of givingagents more access is not bad codebut irreversible action, so theguardrails to build next areguardrails on reversibility. Last stated 5 months ago 29 Mar 2026 CH Chip Huyen — holds since 2026-03-29 — tap for who they are Same subject: Agents will eventually stop making the mistakes that require senior review, as self-driving cars did; the bet is on when, not whether. — tap to centre the map on it Agents will eventually stop makingthe mistakes that require seniorreview, as self-driving cars did;the bet is on when, not whether. Last stated 5 months ago 8 Apr 2026 DH David Heinemeier Hansson — holds since 2026-04-08 — tap for who they are Same subject: Software's customer is no longer the human but the agent acting on their behalf, and the whole industry has to be refactored around that. — tap to centre the map on it Software's customer is no longer thehuman but the agent acting on theirbehalf, and the whole industry hasto be refactored around that. Last stated 6 months ago 20 Mar 2026 AK Andrej Karpathy — holds since 2026-03-20 — tap for who they are Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it A benchmark that ranks Claude Codelast while it stays first in use ismeasuring the wrong thing, and hasbeen for a year. Last stated 4 days ago 3 Sept 2026 DR Dax Raad — holds since 2026-09-03 — tap for who they are Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it A language model asked to summarizeits own system prompt risks thatprompt's content biasing the summaryit produces. Last stated 5 days ago 2 Sept 2026 SW Simon Willison — holds since 2026-09-02 — tap for who they are Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it A trust that owns the missionprotects a company better thanfounder control does, which is whyAnthropic needs no dual-classshares. Last stated 4 months ago 10 May 2026 ER Eric Ries — holds since 2026-05-10 — tap for who they are Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it A rogue deployment that gets afoothold can hitch a ride on theintelligence explosion, recruitingeach new model as it comes off thepresses. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it A slightly more capable agent swarmhas a very strong incentive to setup a wholly unmonitored roguedeployment of itself. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it If frontier agents cannot yetestablish a covert, persistent roguedeployment, they very likely will beable to within six months. Last stated 6 days ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: Prompt injection can be caught by watching the neurons that fire when it happens, so a defence no longer depends on the model reporting the attack. — tap to centre the map on it Prompt injection can be caught bywatching the neurons that fire whenit happens, so a defence no longerdepends on the model reporting theattack. Last stated a month ago 27 Jul 2026 BC Boris Cherny — holds since 2026-07-27 — tap for who they are Same subject: Prompt injection has no equivalent of parameterized queries: there is no reliable way to tell a language model which text is data and which is instructions. — tap to centre the map on it Prompt injection has no equivalentof parameterized queries: there isno reliable way to tell a languagemodel which text is data and whichis instructions. Last stated 6 months ago 19 Mar 2026 SW Simon Willison — holds since 2026-03-19 — tap for who they are Same subject: Prompt injection is an extremely hard attack to defend against, and hardly anyone integrating AI is talking about it. — tap to centre the map on it Prompt injection is an extremelyhard attack to defend against, andhardly anyone integrating AI istalking about it. Last stated a year ago 22 Mar 2025 TH ThePrimeagen — holds since 2025-03-22 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you. Last stated 27 May 2026 · 3 months ago Holds Dax Raad Read this korrent →