korrents

← AI labs tend to downplay safety risks in their public rhetoric even while acting responsibly in practice.

On the map

8 connected korrents · 8 moments on record from 17 Feb 2023 to 4 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject AI labs tend to downplaysafety risks in theirpublic rhetoric even while… ZM Zvi Mowshowitz — holds since 2026-09-04 A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. A technique that letsan AI model's reasoning… ZM Zvi Mowshowitz — holds since 2026-09-03 On AI the state's job does not stop at safety; leaving everything else to the market is a mistake. On AI the state's jobdoes not stop at… KS Keir Starmer — holds since 2025-01-13 Confidence in safety, not capability, will increasingly set the pace of AI progress. Confidence in safety,not capability, will… SA Sam Altman — holds since 2026-08-18 We do not control neural networks well enough to guarantee an AI will not harm humans, so we should be careful about what capabilities we give them. We do not controlneural networks well… AR Armin Ronacher — holds since 2023-02-17 Even well-informed AI insiders frequently develop something like "AI psychosis" from engaging with these ideas. Even well-informed AIinsiders frequently… TC Tyler Cowen — holds since 2026-08-31 Relying on gradual, continuous shifts in AI training behavior to catch misalignment will eventually fail because the dangerous shift itself may be discontinuous. Relying on gradual,continuous shifts in AI… ZM Zvi Mowshowitz — holds since 2026-09-02 Current prompt-injection defenses for AI agents (such as auto mode) are now reliable enough that agents can practically be assumed safe from successful injection attacks. Currentprompt-injection… ZM Zvi Mowshowitz — holds since 2026-09-04 AI-related cybersecurity damage over the next year or two will fall well short of the harm caused by Covid or global warming AI-relatedcybersecurity damage… TC Tyler Cowen — holds since 2026-09-01
same subject