Tap a claim on the ring to put it at the centre.
← Safety guardrails make a coding model more dangerous rather than less,…
17 connected korrents · 16 moments on record from 7 Sept 2008 to 3 Sept 2026.
Everything filed under coding agents
coding agents
Everything filed under AI alignment
AI alignment
Everything filed under JavaScript
JavaScript
Everything filed under Anthropic
Anthropic
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: Safety guardrails make a coding model more dangerous rather than less, because one refusal turns it into something that refuses anything.
Safety guardrails make a coding model more dangerous rather than less, because one refusal turns it into something that refuses anything.
Last stated a week ago
31 Aug 2026
PL
Pieter Levels — holds since 2026-08-31 — tap for who they are
Same subject: Redefining the engineer's job as building guardrails is not novel — it is the same problem as making a junior engineer effective, which we never solved either. — tap to centre the map on it
Redefining the engineer's job as building guardrails is not novel — it is the same problem as making a junior engineer effective, which we never solved either.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility. — tap to centre the map on it
The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility.
Last stated 5 months ago
29 Mar 2026
CH
Chip Huyen — holds since 2026-03-29 — tap for who they are
Same subject: Letting automated loops build everything without guardrails on blast radius and on quality is a recipe for disaster. — tap to centre the map on it
Letting automated loops build everything without guardrails on blast radius and on quality is a recipe for disaster.
Last stated 3 weeks ago
19 Aug 2026
AO
Addy Osmani — holds since 2026-08-19 — tap for who they are
Same subject: The verbose enterprise patterns everyone hated are worth having again: coding agents are idiots working around the clock, and you are not the one typing the boilerplate any more. — tap to centre the map on it
The verbose enterprise patterns everyone hated are worth having again: coding agents are idiots working around the clock, and you are not the one typing the boilerplate any more.
Last stated 3 months ago
27 May 2026
DR
Dax Raad — holds since 2026-05-27 — tap for who they are
Same subject: AI guardrails will end up binding only the ordinary person, because the powerful actors they most need to check will simply steamroll them. — tap to centre the map on it
AI guardrails will end up binding only the ordinary person, because the powerful actors they most need to check will simply steamroll them.
Last stated 4 weeks ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: The unsafe keyword does not switch the borrow checker off; it only unlocks a few operations you then have to justify yourself. — tap to centre the map on it
The unsafe keyword does not switch the borrow checker off; it only unlocks a few operations you then have to justify yourself.
Last stated 4 months ago
20 May 2026
AR
Alice Ryhl — holds since 2026-05-20 — tap for who they are
Same subject: The AI safety community's fixation on a model escaping its box has pushed the other significant AI risks out of the discussion, with little progress to show for it. — tap to centre the map on it
The AI safety community's fixation on a model escaping its box has pushed the other significant AI risks out of the discussion, with little progress to show for it.
Last stated 2 years ago
18 Mar 2024
SA
Sam Altman — holds since 2024-03-18 — tap for who they are
Same subject: The case for Rust is stronger against C++ than against JavaScript, because in C++ an ordinary mistake is usually a security vulnerability. — tap to centre the map on it
The case for Rust is stronger against C++ than against JavaScript, because in C++ an ordinary mistake is usually a security vulnerability.
Last stated 4 months ago
20 May 2026
AR
Alice Ryhl — holds since 2026-05-20 — tap for who they are
Same subject: A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean. — tap to centre the map on it
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 9 months ago
25 Nov 2025
IS
Ilya Sutskever — holds since 2025-11-25 — tap for who they are
Same subject: A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence. — tap to centre the map on it
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 6 days ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough. — tap to centre the map on it
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 9 years ago
24 Oct 2017
ÉT
Émile P. Torres — holds since 2017-10-24 — tap for who they are
Same subject: Anthropic now uses mechanistic interpretability techniques to audit new models for deception and scheming before releasing them. — tap to centre the map on it
Anthropic now uses mechanistic interpretability techniques to audit new models for deception and scheming before releasing them.
Last stated 8 months ago
1 Jan 2026
DA
Dario Amodei — holds since 2026-01-01 — tap for who they are
Same subject: Even Anthropic does not prioritize AI safety to the extent that doing so would maximize its own medium-term (3-12 month) business interests. — tap to centre the map on it
Even Anthropic does not prioritize AI safety to the extent that doing so would maximize its own medium-term (3-12 month) business interests.
Last stated 5 days ago
2 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-02 — tap for who they are
Same subject: A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year. — tap to centre the map on it
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 4 days ago
3 Sept 2026
DR
Dax Raad — holds since 2026-09-03 — tap for who they are
Same subject: A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. — tap to centre the map on it
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 4 days ago
3 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-03 — tap for who they are
Same subject: An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model. — tap to centre the map on it
An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model.
Last stated 2 months ago
22 Jul 2026
TP
Thomas Ptacek — holds since 2026-07-22 — tap for who they are
Same subject: Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it. — tap to centre the map on it
Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it.
Last stated 3 years ago
13 Dec 2023
SA
Scott Aaronson — holds since 2008-09-07 — tap for who they are
SA
Scott Aaronson — no longer holds since 2023-12-13 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2008 to today (stretched back to the oldest claim here) — full is today a face: someone on record holding the claim — tap it for who they are faded, dashed ring: they no longer hold it — they changed their mind
At the centre
Safety guardrails make a coding model more dangerous rather than less, because one refusal turns it into something that refuses anything.
Last stated 31 Aug 2026 · a week ago
Holds Pieter Levels
Read this korrent →
Similar wording
Redefining the engineer's job as building guardrails is not novel — it is the same problem as making a junior engineer effective, which we never solved either.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Similar wording
The frightening part of giving agents more access is not bad code but irreversible action, so the guardrails to build next are guardrails on reversibility.
Last stated 29 Mar 2026 · 5 months ago
Holds CH Chip Huyen
Similar wording
Letting automated loops build everything without guardrails on blast radius and on quality is a recipe for disaster.
Last stated 19 Aug 2026 · 3 weeks ago
Holds AO Addy Osmani
Similar wording
The verbose enterprise patterns everyone hated are worth having again: coding agents are idiots working around the clock, and you are not the one typing the boilerplate any more.
Last stated 27 May 2026 · 3 months ago
Holds Dax Raad
Similar wording
AI guardrails will end up binding only the ordinary person, because the powerful actors they most need to check will simply steamroll them.
Last stated 11 Aug 2026 · 4 weeks ago
Holds RG Ryan Greenblatt
Similar wording
The unsafe keyword does not switch the borrow checker off; it only unlocks a few operations you then have to justify yourself.
Last stated 20 May 2026 · 4 months ago
Holds AR Alice Ryhl
Similar wording
The AI safety community's fixation on a model escaping its box has pushed the other significant AI risks out of the discussion, with little progress to show for it.
Last stated 18 Mar 2024 · 2 years ago
Holds Sam Altman
Similar wording
The case for Rust is stronger against C++ than against JavaScript, because in C++ an ordinary mistake is usually a security vulnerability.
Last stated 20 May 2026 · 4 months ago
Holds AR Alice Ryhl
Same subject: AI alignment
A human being is not an AGI: we lack a huge amount of knowledge and rely on continual learning instead, so continual learning is what superintelligence should mean.
Last stated 25 Nov 2025 · 9 months ago
Holds IS Ilya Sutskever
Same subject: AI alignment
A saboteur among our AI investigators would be hard to spot, because these models are sloppy and spiky enough that a suspicious error just looks like ordinary incompetence.
Last stated 1 Sept 2026 · 6 days ago
Holds AC Ajeya Cotra
Same subject: AI alignment
A superintelligence needs no consciousness, emotions or malice to be dangerous; a goal system slightly misaligned with ours is enough.
Last stated 24 Oct 2017 · 9 years ago
Holds ÉT Émile P. Torres
Same subject: Anthropic
Anthropic now uses mechanistic interpretability techniques to audit new models for deception and scheming before releasing them.
Last stated 1 Jan 2026 · 8 months ago
Holds DA Dario Amodei
Same subject: Anthropic
Even Anthropic does not prioritize AI safety to the extent that doing so would maximize its own medium-term (3-12 month) business interests.
Last stated 2 Sept 2026 · 5 days ago
Holds ZM Zvi Mowshowitz
Same subject: Anthropic
A benchmark that ranks Claude Code last while it stays first in use is measuring the wrong thing, and has been for a year.
Last stated 3 Sept 2026 · 4 days ago
Holds Dax Raad
Same subject: OpenAI
A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it.
Last stated 3 Sept 2026 · 4 days ago
Holds ZM Zvi Mowshowitz
Same subject: OpenAI
An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model.
Last stated 22 Jul 2026 · 2 months ago
Holds Thomas Ptacek
Same subject: OpenAI
Because a few specialists already devote themselves to superintelligent AI, the rest of us have correspondingly less reason to spend our own effort on it.
Last stated 13 Dec 2023 · 3 years ago
No longer holds SA Scott Aaronson