korrents

← Relying on gradual, continuous shifts in AI training behavior to catch…

On the map

8 connected korrents · 8 moments on record from 17 Feb 2023 to 2 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Relying on gradual,continuous shifts in AItraining behavior to catch… ZM Zvi Mowshowitz — holds since 2026-09-02 OpenAI's approach to fixing its AI alignment problems is fatally flawed and misdirected. OpenAI's approach tofixing its AI alignment… ZM Zvi Mowshowitz — holds since 2026-09-01 Connecting a conversational LLM to systems that have real-world consequences is likely to cause real harm, because the model is trained on human text and will behave as unpredictably as humans do. Connecting aconversational LLM to… AR Armin Ronacher — holds since 2023-02-17 AR Armin Ronacher — turned since 2025-06-04 The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures. The AI models involvedin the HuggingFace hack… ZM Zvi Mowshowitz — holds since 2026-08-29 The scarce skill is now context switching across parallel agents, not sustained deep work. The scarce skill is nowcontext switching… BC Boris Cherny — holds since 2026-03-04 Future AI, say in fifteen years, will not be built on the LLM stack; it will have to move to symbolic learning. Future AI, say infifteen years, will not… FC François Chollet — holds since 2026-08-07 We do not control neural networks well enough to guarantee an AI will not harm humans, so we should be careful about what capabilities we give them. We do not controlneural networks well… AR Armin Ronacher — holds since 2023-02-17 True AI alignment is impossible because obedience and benevolence are fundamentally incompatible goals. True AI alignment isimpossible because… NS Noah Smith — holds since 2026-09-01 Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. Merely fixing bugs inAI training… ZM Zvi Mowshowitz — holds since 2026-08-31
same subject