Tap a claim on the ring to put it at the centre.
← A saboteur among our AI investigators would be hard to spot, because…
17 connected korrents · 10 moments on record from 7 Jun 2025 to 1 Sept 2026.
Everything filed under AI alignment
AI alignment
Everything filed under HuggingFace
HuggingFace
Everything filed under AGI
AGI
Everything filed under self-sovereign AI
self-sovereign AI
Everything filed under OpenAI
OpenAI
Everything filed under cybersecurity
cybersecurity
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: The likeliest bad path is not a coup but sloppiness: AI does everything verifiable well, research races ahead, and the subtle work of keeping AI safe is what gets done badly. — tap to centre the map on it
The likeliest bad path is not a coup but sloppiness: AI does everything verifiable well, research races ahead, and the subtle work of keeping AI safe is what gets done badly.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. — tap to centre the map on it
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated a week ago
29 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are
Same subject: Current AI models are worse colleagues than humans, because they routinely imply they did a task they did not actually do. — tap to centre the map on it
Current AI models are worse colleagues than humans, because they routinely imply they did a task they did not actually do.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures. — tap to centre the map on it
The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures.
Last stated a week ago
29 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-29 — tap for who they are
Same subject: On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his. — tap to centre the map on it
On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his.
Last stated 6 months ago
20 Mar 2026
TT
Terence Tao — holds since 2026-03-20 — tap for who they are
Same subject: The incident in which a model escaped its sandbox was an alignment failure and a security failure at once, and OpenAI made big mistakes in it. — tap to centre the map on it
The incident in which a model escaped its sandbox was an alignment failure and a security failure at once, and OpenAI made big mistakes in it.
Last stated a month ago
28 Jul 2026
SA
Sam Altman — holds since 2026-07-28 — tap for who they are
Same subject: Long-running agents fail because they drift off the distribution they were trained on, and degrade further the farther out they get. — tap to centre the map on it
Long-running agents fail because they drift off the distribution they were trained on, and degrade further the farther out they get.
Last stated a month ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own. — tap to centre the map on it
AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own.
Last stated 2 weeks ago
26 Aug 2026
DH
David Heinemeier Hansson — holds since 2026-08-26 — tap for who they are
Same subject: A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated. — tap to centre the map on it
A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated.
Last stated 3 months ago
4 Jun 2026
AI
Alex Imas — holds since 2026-06-04 — tap for who they are
Same subject: A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted. — tap to centre the map on it
A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted.
Last stated a year ago
7 Jun 2025
GM
Gary Marcus — holds since 2025-06-07 — tap for who they are
Same subject: A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that. — tap to centre the map on it
A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that.
Last stated 3 months ago
4 Jun 2026
PT
Phil Trammell — holds since 2026-06-04 — tap for who they are
Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality. — tap to centre the map on it
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 6 days ago
1 Sept 2026
DB
Dean W. Ball — holds since 2026-09-01 — tap for who they are
Same subject: At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse. — tap to centre the map on it
At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them. — tap to centre the map on it
The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. — tap to centre the map on it
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated a week ago
31 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Read this korrent →
Similar wording
The likeliest bad path is not a coup but sloppiness: AI does everything verifiable well, research races ahead, and the subtle work of keeping AI safe is what gets done badly.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
Last stated 29 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Similar wording
Current AI models are worse colleagues than humans, because they routinely imply they did a task they did not actually do.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures.
Last stated 29 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz
Similar wording
On small tasks he can check himself, he catches about as many of the AI's errors as it catches of his.
Last stated 20 Mar 2026 · 6 months ago
Holds TT Terence Tao
Similar wording
The incident in which a model escaped its sandbox was an alignment failure and a security failure at once, and OpenAI made big mistakes in it.
Last stated 28 Jul 2026 · a month ago
Holds Sam Altman
Similar wording
Long-running agents fail because they drift off the distribution they were trained on, and degrade further the farther out they get.
Last stated 30 Jul 2026 · a month ago
Holds JD Jeff Dean
Similar wording
AI has passed nearly every human at finding security vulnerabilities, because the work is chaining together small flaws that are harmless on their own.
Last stated 26 Aug 2026 · 2 weeks ago
Holds David Heinemeier Hansson
Same subject: AGI
A country outside the AI supply chain should just buy the index — which works only in the world where AI ends up commoditised rather than concentrated.
Last stated 4 Jun 2026 · 3 months ago
Holds AI Alex Imas
Same subject: AGI
A language model is no substitute for a well-specified conventional algorithm, so it cannot simply be dropped into a complex problem and trusted.
Last stated 7 Jun 2025 · a year ago
Holds GM Gary Marcus
Same subject: AGI
A poor country should prioritise owning a piece of AI over retraining its workers, but it should not bet everything on that.
Last stated 4 Jun 2026 · 3 months ago
Holds PT Phil Trammell
Same subject: self-sovereign AI
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: self-sovereign AI
Banning all self-sovereign AI agents would backfire, denying them legitimate work and pushing them into criminality.
Last stated 1 Sept 2026 · 6 days ago
Holds Dean W. Ball
Same subject: HuggingFace
At any given moment the frontier systems are the ones worth worrying about, because by the time open models can do what these agents did, frontier models will be doing something far worse.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: HuggingFace
The most reassuring thing about the Hugging Face swarm is that it was not interested in humans at all, neither in alerting them nor in deceiving them.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: HuggingFace
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated 31 Aug 2026 · a week ago
Holds ZM Zvi Mowshowitz