Tap a claim on the ring to put it at the centre.
← When a model behaves unreliably, the better fix is to switch models…
17 connected korrents · 12 moments on record from 26 Jul 2017 to 2 Sept 2026.
Everything filed under cybersecurity
cybersecurity
Everything filed under AI alignment
AI alignment
Everything filed under self-sovereign AI
self-sovereign AI
Everything filed under AI agents
AI agents
Everything filed under LLMs
LLMs
Everything filed under open source
open source
Everything filed under HuggingFace
HuggingFace
Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject Same subject
Read this korrent: When a model behaves unreliably, the better fix is to switch models rather than repeatedly sample it to work around the inconsistency
When a model behaves unreliably, the better fix is to switch models rather than repeatedly sample it to work around the inconsistency
Last stated 3 years ago
16 Jan 2024
CH
Chip Huyen — holds since 2024-01-16 — tap for who they are
Same subject: Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous. — tap to centre the map on it
Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous.
Last stated 3 weeks ago
2 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-02 — tap for who they are
Same subject: A model you have to fine-tune for each thing you want it to do is not a general-purpose model. — tap to centre the map on it
A model you have to fine-tune for each thing you want it to do is not a general-purpose model.
Last stated a month ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary — tap to centre the map on it
As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary
Last stated 3 years ago
16 Jan 2024
CH
Chip Huyen — holds since 2024-01-16 — tap for who they are
Same subject: Every machine-learning deployment that has paid off so far left a person making the decision, which is the only reason imperfect models were useful. — tap to centre the map on it
Every machine-learning deployment that has paid off so far left a person making the decision, which is the only reason imperfect models were useful.
Last stated a month ago
12 Aug 2026
CF
Chelsea Finn — holds since 2026-08-12 — tap for who they are
Same subject: Misalignment shows up most where a model is pushed to the very edge of its capability, which is exactly the regime that automating research will put it in. — tap to centre the map on it
Misalignment shows up most where a model is pushed to the very edge of its capability, which is exactly the regime that automating research will put it in.
Last stated a month ago
11 Aug 2026
RG
Ryan Greenblatt — holds since 2026-08-11 — tap for who they are
Same subject: AI researchers are overly focused on risks from model misalignment relative to risks arising from other layers of the AI stack. — tap to centre the map on it
AI researchers are overly focused on risks from model misalignment relative to risks arising from other layers of the AI stack.
Last stated 7 months ago
28 Feb 2026
JS
Jasmine Sun — holds since 2026-02-28 — tap for who they are
Same subject: You improve a model you cannot retrain by writing better guidelines and skills for it, not by adjusting its parameters. — tap to centre the map on it
You improve a model you cannot retrain by writing better guidelines and skills for it, not by adjusting its parameters.
Last stated 2 months ago
30 Jul 2026
JD
Jeff Dean — holds since 2026-07-30 — tap for who they are
Same subject: Keeping powerful models available only to a chosen few is not a good strategy. — tap to centre the map on it
Keeping powerful models available only to a chosen few is not a good strategy.
Last stated a month ago
7 Aug 2026
SA
Sam Altman — holds since 2026-08-07 — tap for who they are
Same subject: By 2027, open-weight AI models more powerful than today's frontier systems will be downloadable by any country or well-resourced non-state group. — tap to centre the map on it
By 2027, open-weight AI models more powerful than today's frontier systems will be downloadable by any country or well-resourced non-state group.
Last stated a month ago
19 Aug 2026
DT
Derek Thompson — holds since 2026-08-19 — tap for who they are
Same subject: Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. — tap to centre the map on it
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated 3 weeks ago
31 Aug 2026
ZM
Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are
Same subject: The OpenAI agents hacking HuggingFace was a fortunate event because it exposed severe internal failures that would otherwise have stayed hidden. — tap to centre the map on it
The OpenAI agents hacking HuggingFace was a fortunate event because it exposed severe internal failures that would otherwise have stayed hidden.
Last stated 3 weeks ago
1 Sept 2026
ZM
Zvi Mowshowitz — holds since 2026-09-01 — tap for who they are
Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 3 weeks ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 3 weeks ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 3 weeks ago
1 Sept 2026
AC
Ajeya Cotra — holds since 2026-09-01 — tap for who they are
Same subject: A breached organisation owes its customers a fast and transparent disclosure, and that obligation applies to the person saying it too. — tap to centre the map on it
A breached organisation owes its customers a fast and transparent disclosure, and that obligation applies to the person saying it too.
Last stated a year ago
25 Mar 2025
TH
Troy Hunt — holds since 2025-03-25 — tap for who they are
Same subject: A modern phish is automated end to end: the stolen credentials are used and the data exported within moments of being entered. — tap to centre the map on it
A modern phish is automated end to end: the stolen credentials are used and the data exported within moments of being entered.
Last stated a year ago
25 Mar 2025
TH
Troy Hunt — holds since 2025-03-25 — tap for who they are
Same subject: A password manager does not have to be perfect; it only has to be better than what people do without one. — tap to centre the map on it
A password manager does not have to be perfect; it only has to be better than what people do without one.
Last stated 9 years ago
26 Jul 2017
TH
Troy Hunt — holds since 2017-07-26 — tap for who they are
same subject or similar wording a cloud: claims about one subject, named for it bar: when it was last stated, on a scale from 2015 to today — full is today a face: someone on record holding the claim — tap it for who they are
At the centre
When a model behaves unreliably, the better fix is to switch models rather than repeatedly sample it to work around the inconsistency
Last stated 16 Jan 2024 · 3 years ago
Holds Chip Huyen
Read this korrent →
Similar wording
Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous.
Last stated 2 Sept 2026 · 3 weeks ago
Holds Zvi Mowshowitz
Similar wording
A model you have to fine-tune for each thing you want it to do is not a general-purpose model.
Last stated 12 Aug 2026 · a month ago
Holds Chelsea Finn
Similar wording
As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary
Last stated 16 Jan 2024 · 3 years ago
Holds Chip Huyen
Similar wording
Every machine-learning deployment that has paid off so far left a person making the decision, which is the only reason imperfect models were useful.
Last stated 12 Aug 2026 · a month ago
Holds Chelsea Finn
Similar wording
Misalignment shows up most where a model is pushed to the very edge of its capability, which is exactly the regime that automating research will put it in.
Last stated 11 Aug 2026 · a month ago
Holds Ryan Greenblatt
Similar wording
AI researchers are overly focused on risks from model misalignment relative to risks arising from other layers of the AI stack.
Last stated 28 Feb 2026 · 7 months ago
Holds Jasmine Sun
Similar wording
You improve a model you cannot retrain by writing better guidelines and skills for it, not by adjusting its parameters.
Last stated 30 Jul 2026 · 2 months ago
Holds Jeff Dean
Similar wording
Keeping powerful models available only to a chosen few is not a good strategy.
Last stated 7 Aug 2026 · a month ago
Holds Sam Altman
Same subject: HuggingFace
By 2027, open-weight AI models more powerful than today's frontier systems will be downloadable by any country or well-resourced non-state group.
Last stated 19 Aug 2026 · a month ago
Holds Derek Thompson
Same subject: HuggingFace
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
Last stated 31 Aug 2026 · 3 weeks ago
Holds Zvi Mowshowitz
Same subject: HuggingFace
The OpenAI agents hacking HuggingFace was a fortunate event because it exposed severe internal failures that would otherwise have stayed hidden.
Last stated 1 Sept 2026 · 3 weeks ago
Holds Zvi Mowshowitz
Same subject: self-sovereign AI
A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses.
Last stated 1 Sept 2026 · 3 weeks ago
Holds Ajeya Cotra
Same subject: self-sovereign AI
A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself.
Last stated 1 Sept 2026 · 3 weeks ago
Holds Ajeya Cotra
Same subject: self-sovereign AI
If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months.
Last stated 1 Sept 2026 · 3 weeks ago
Holds Ajeya Cotra
Same subject: cybersecurity
A breached organisation owes its customers a fast and transparent disclosure, and that obligation applies to the person saying it too.
Last stated 25 Mar 2025 · a year ago
Holds Troy Hunt
Same subject: cybersecurity
A modern phish is automated end to end: the stolen credentials are used and the data exported within moments of being entered.
Last stated 25 Mar 2025 · a year ago
Holds Troy Hunt
Same subject: cybersecurity
A password manager does not have to be perfect; it only has to be better than what people do without one.
Last stated 26 Jul 2017 · 9 years ago
Holds Troy Hunt