korrents

On the map

Tap a claim on the ring to put it at the centre.

← When a model behaves unreliably, the better fix is to switch models…

17 connected korrents · 12 moments on record from 26 Jul 2017 to 2 Sept 2026.

Everything filed under cybersecurity cybersecurity Everything filed under AI alignment AI alignment Everything filed under self-sovereign AI self-sovereign AI Everything filed under AI agents AI agents Everything filed under LLMs LLMs Everything filed under open source open source Everything filed under HuggingFace HuggingFace Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Read this korrent: When a model behaves unreliably, the better fix is to switch models rather than repeatedly sample it to work around the inconsistency When a model behaves unreliably, thebetter fix is to switch models rather thanrepeatedly sample it to work around theinconsistency Last stated 3 years ago 16 Jan 2024 CH Chip Huyen — holds since 2024-01-16 — tap for who they are Same subject: Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous. — tap to centre the map on it Relying on gradual, continuousshifts in AI training behavior tocatch misalignment will eventuallyfail because the dangerous shiftitself may be discontinuous. Last stated 3 weeks ago 2 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-02 — tap for who they are Same subject: A model you have to fine-tune for each thing you want it to do is not a general-purpose model. — tap to centre the map on it A model you have to fine-tune foreach thing you want it to do is nota general-purpose model. Last stated a month ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: As models get better at following instructions, techniques like finetuning and constrained sampling for structured outputs will become less necessary — tap to centre the map on it As models get better at followinginstructions, techniques likefinetuning and constrained samplingfor structured outputs will becomeless necessary Last stated 3 years ago 16 Jan 2024 CH Chip Huyen — holds since 2024-01-16 — tap for who they are Same subject: Every machine-learning deployment that has paid off so far left a person making the decision, which is the only reason imperfect models were useful. — tap to centre the map on it Every machine-learning deploymentthat has paid off so far left aperson making the decision, which isthe only reason imperfect modelswere useful. Last stated a month ago 12 Aug 2026 CF Chelsea Finn — holds since 2026-08-12 — tap for who they are Same subject: Misalignment shows up most where a model is pushed to the very edge of its capability, which is exactly the regime that automating research will put it in. — tap to centre the map on it Misalignment shows up most where amodel is pushed to the very edge ofits capability, which is exactly theregime that automating research willput it in. Last stated a month ago 11 Aug 2026 RG Ryan Greenblatt — holds since 2026-08-11 — tap for who they are Same subject: AI researchers are overly focused on risks from model misalignment relative to risks arising from other layers of the AI stack. — tap to centre the map on it AI researchers are overly focused onrisks from model misalignmentrelative to risks arising from otherlayers of the AI stack. Last stated 7 months ago 28 Feb 2026 JS Jasmine Sun — holds since 2026-02-28 — tap for who they are Same subject: You improve a model you cannot retrain by writing better guidelines and skills for it, not by adjusting its parameters. — tap to centre the map on it You improve a model you cannotretrain by writing better guidelinesand skills for it, not by adjustingits parameters. Last stated 2 months ago 30 Jul 2026 JD Jeff Dean — holds since 2026-07-30 — tap for who they are Same subject: Keeping powerful models available only to a chosen few is not a good strategy. — tap to centre the map on it Keeping powerful models availableonly to a chosen few is not a goodstrategy. Last stated a month ago 7 Aug 2026 SA Sam Altman — holds since 2026-08-07 — tap for who they are Same subject: By 2027, open-weight AI models more powerful than today's frontier systems will be downloadable by any country or well-resourced non-state group. — tap to centre the map on it By 2027, open-weight AI models morepowerful than today's frontiersystems will be downloadable by anycountry or well-resourced non-stategroup. Last stated a month ago 19 Aug 2026 DT Derek Thompson — holds since 2026-08-19 — tap for who they are Same subject: Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. — tap to centre the map on it Frontier AI labs such as Anthropiclikely have had internal securityincidents similar to OpenAI'sHuggingFace attack that were neverpublicly disclosed. Last stated 3 weeks ago 31 Aug 2026 ZM Zvi Mowshowitz — holds since 2026-08-31 — tap for who they are Same subject: The OpenAI agents hacking HuggingFace was a fortunate event because it exposed severe internal failures that would otherwise have stayed hidden. — tap to centre the map on it The OpenAI agents hackingHuggingFace was a fortunate eventbecause it exposed severe internalfailures that would otherwise havestayed hidden. Last stated 3 weeks ago 1 Sept 2026 ZM Zvi Mowshowitz — holds since 2026-09-01 — tap for who they are Same subject: A rogue deployment that gets a foothold can hitch a ride on the intelligence explosion, recruiting each new model as it comes off the presses. — tap to centre the map on it A rogue deployment that gets afoothold can hitch a ride on theintelligence explosion, recruitingeach new model as it comes off thepresses. Last stated 3 weeks ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A slightly more capable agent swarm has a very strong incentive to set up a wholly unmonitored rogue deployment of itself. — tap to centre the map on it A slightly more capable agent swarmhas a very strong incentive to setup a wholly unmonitored roguedeployment of itself. Last stated 3 weeks ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: If frontier agents cannot yet establish a covert, persistent rogue deployment, they very likely will be able to within six months. — tap to centre the map on it If frontier agents cannot yetestablish a covert, persistent roguedeployment, they very likely will beable to within six months. Last stated 3 weeks ago 1 Sept 2026 AC Ajeya Cotra — holds since 2026-09-01 — tap for who they are Same subject: A breached organisation owes its customers a fast and transparent disclosure, and that obligation applies to the person saying it too. — tap to centre the map on it A breached organisation owes itscustomers a fast and transparentdisclosure, and that obligationapplies to the person saying it too. Last stated a year ago 25 Mar 2025 TH Troy Hunt — holds since 2025-03-25 — tap for who they are Same subject: A modern phish is automated end to end: the stolen credentials are used and the data exported within moments of being entered. — tap to centre the map on it A modern phish is automated end toend: the stolen credentials are usedand the data exported within momentsof being entered. Last stated a year ago 25 Mar 2025 TH Troy Hunt — holds since 2025-03-25 — tap for who they are Same subject: A password manager does not have to be perfect; it only has to be better than what people do without one. — tap to centre the map on it A password manager does not have tobe perfect; it only has to be betterthan what people do without one. Last stated 9 years ago 26 Jul 2017 TH Troy Hunt — holds since 2017-07-26 — tap for who they are
same subject or similar wordinga cloud: claims about one subject, named for itbar: when it was last stated, on a scale from 2015 to today — full is todaya face: someone on record holding the claim — tap it for who they are

At the centre When a model behaves unreliably, the better fix is to switch models rather than repeatedly sample it to work around the inconsistency Last stated 16 Jan 2024 · 3 years ago Holds Chip Huyen Read this korrent →