Tap a claim on the ring to put it at the centre.
← The judgement call about whether to absorb a hack is distorted right…
17 connected korrents · 13 moments on record from 22 Mar 2025 to 3 Sept 2026.
Everything filed under Anthropic
Anthropic
Everything filed under prompt injection
prompt injection
Everything filed under self-sovereign AI
self-sovereign AI
Everything filed under AI agents
AI agents
Everything filed under coding agents
coding agents
Everything filed under AI alignment
AI alignment
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you.
The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart. — tap to centre the map on it
The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat. — tap to centre the map on it
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated a week ago
29 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are
Same subject: The prickle you used to feel writing a hack is muted now, so the landmines are still there but the feedback loop that kept your judgement honest is gone. — tap to centre the map on it
The prickle you used to feel writing a hack is muted now, so the landmines are still there but the feedback loop that kept your judgement honest is gone.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: You have to keep enough understanding of how a system works to fix it yourself, because the alternative is hoping and praying the agent can figure it out. — tap to centre the map on it
You have to keep enough understanding of how a system works to fix it yourself, because the alternative is hoping and praying the agent can figure it out.
Last stated 3 weeks ago
19 Aug 2026
AO
Addy Osmani — holds since 2026-08-19 — tap for who they are
Same subject: The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility. — tap to centre the map on it
The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility.
Last stated 5 months ago
29 Mar 2026
CH
Chip Huyen — holds since 2026-03-29 — tap for who they are
Same subject: Agents will eventually stop making the mistakes that require senior review, as self-driving cars did; the bet is on when, not whether. — tap to centre the map on it
Agents will eventually stop making the mistakes that require senior review, as self-driving cars did; the bet is on when, not whether.
Last stated 5 months ago
8 Apr 2026
DH
David Heinemeier Hansson — holds since 2026-04-08 — tap for who they are
Same subject: Software's customer is no longer the human but the agent acting on their behalf, and the whole industry has to be refactored around that. — tap to centre the map on it
Software's customer is no longer the human but the agent acting on their behalf, and the whole industry has to be refactored around that.
Last stated 6 months ago
20 Mar 2026
AK
Andrej Karpathy — holds since 2026-03-20 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces. — tap to centre the map on it
A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces.
Last stated 5 days ago
2 Sept 2026
SW
Simon Willison — holds since 2026-09-02 — tap for who they are
Same subject: A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares. — tap to centre the map on it
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 4 months ago
10 May 2026
ER
Eric Ries — holds since 2026-05-10 — tap for who they are
Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Prompt injection can be caught by watching the neurons that fire when it happens, so a defence no longer depends on the model reporting the attack. — tap to centre the map on it
Prompt injection can be caught by watching the neurons that fire when it happens, so a defence no longer depends on the model reporting the attack.
Last stated a month ago
27 Jul 2026
BC
Boris Cherny — holds since 2026-07-27 — tap for who they are
Same subject: Prompt injection has no equivalent of parameterized queries: there is no reliable way to tell a language model which text is data and which is instructions. — tap to centre the map on it
Prompt injection has no equivalent of parameterized queries: there is no reliable way to tell a language model which text is data and which is instructions.
Last stated 6 months ago
19 Mar 2026
SW
Simon Willison — holds since 2026-03-19 — tap for who they are
Same subject: Prompt injection is an extremely hard attack to defend against, and hardly anyone integrating AI is talking about it. — tap to centre the map on it
Prompt injection is an extremely hard attack to defend against, and hardly anyone integrating AI is talking about it.
Last stated a year ago
22 Mar 2025
TH
ThePrimeagen — holds since 2025-03-22 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
The judgement call about whether to absorb a hack is distorted right now, because the agent will simply do the hacky thing and deal with the consequences for you.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Read this korrent →
Similar wording
The winning move in coding agents was inverted: take share with a merely good-enough harness first, then go back and make the harness smart.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Similar wording
The fix for reward hacking is to take out the environments that reward it, not to add penalties the agent must then balance against the temptation to cheat.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Similar wording
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated 29 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Similar wording
The prickle you used to feel writing a hack is muted now, so the landmines are still there but the feedback loop that kept your judgement honest is gone.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Similar wording
You have to keep enough understanding of how a system works to fix it yourself, because the alternative is hoping and praying the agent can figure it out.
Last stated 19 Aug 2026 · 3 weeks ago
Holds AO Addy Osmani
Similar wording
The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility.
Last stated 29 Mar 2026 · 5 months ago
Holds CH Chip Huyen
Similar wording
Agents will eventually stop making the mistakes that require senior review, as self-driving cars did; the bet is on when, not whether.
Last stated 8 Apr 2026 · 5 months ago
Holds David Heinemeier Hansson
Similar wording
Software's customer is no longer the human but the agent acting on their behalf, and the whole industry has to be refactored around that.
Last stated 20 Mar 2026 · 6 months ago
Holds Andrej Karpathy
Same subject: Anthropic
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: Anthropic
A language model asked to summarize its own system prompt risks that prompt's content biasing the summary it produces.
Last stated 2 Sept 2026 · 5 days ago
Holds Simon Willison
Same subject: Anthropic
A trust that owns the mission protects a company better than founder control does, which is why Anthropic needs no dual-class shares.
Last stated 10 May 2026 · 4 months ago
Holds ER Eric Ries
Same subject: self-sovereign AI
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: prompt injection
Prompt injection can be caught by watching the neurons that fire when it happens, so a defence no longer depends on the model reporting the attack.
Last stated 27 Jul 2026 · a month ago
Holds Boris Cherny
Same subject: prompt injection
Prompt injection has no equivalent of parameterized queries: there is no reliable way to tell a language model which text is data and which is instructions.
Last stated 19 Mar 2026 · 6 months ago
Holds Simon Willison
Same subject: prompt injection
Prompt injection is an extremely hard attack to defend against, and hardly anyone integrating AI is talking about it.
Last stated 22 Mar 2025 · a year ago
Holds TH ThePrimeagen