korrents

← Disclosure of rogue AI agent incidents should be mandatory for AI labs, not left to their own discretion.

On the map

8 connected korrents · 9 moments on record from 12 May 2025 to 6 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Disclosure of rogue AIagent incidents should bemandatory for AI labs, not… ZM Zvi Mowshowitz — holds since 2026-09-06 A technique that lets an AI model's reasoning shift outside its visible Chain of Thought is dangerous, both because it works and because a leading lab is willing to deploy it. A technique that letsan AI model's reasoning… ZM Zvi Mowshowitz — holds since 2026-09-03 A tiny language model inventing a plausible-sounding name is the same phenomenon as a large one confidently stating a false fact. A tiny language modelinventing a… AK Andrej Karpathy — holds since 2026-02-12 AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. AI agents werereasonable to assume a… ZM Zvi Mowshowitz — holds since 2026-08-29 An open weights model from 2025 with a pentest harness could already escape a sandbox and hack most networks; the surprise says more about the sandbox than the model. An open weights modelfrom 2025 with a… TP Thomas Ptacek — holds since 2026-07-22 Chatbot interfaces are the wrong tool for meaningful coding work, because you are mostly hoping the model recalls the right answer and correcting it requires a human in the loop. Chatbot interfaces arethe wrong tool for… MH Mitchell Hashimoto — holds since 2026-02-05 ChatGPT exploded because of its obvious chat interface, not its technology, which had launched earlier in a different wrapper that few noticed. ChatGPT explodedbecause of its obvious… JZ Julie Zhuo — holds since 2025-05-12 Claude Haiku hallucinates badly enough that its use inside Claude Code's WebFetch tool is a hallucination risk on every fetched URL. Claude Haikuhallucinates badly… SW Simon Willison — holds since 2026-08-10 Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. Frontier AI labs suchas Anthropic likely… ZM Zvi Mowshowitz — holds since 2026-08-31
same subject