korrents

← Current prompt-injection defenses for AI agents (such as auto mode)…

On the map

8 connected korrents · 8 moments on record from 17 Feb 2023 to 4 Sept 2026.

Same subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subjectSame subject Current prompt-injectiondefenses for AI agents(such as auto mode) are now… ZM Zvi Mowshowitz — holds since 2026-09-04 Auto-mode does not yet convincingly fix prompt-injection risk for coding agents. Auto-mode does not yetconvincingly fix… SW Simon Willison — holds since 2026-08-08 We do not control neural networks well enough to guarantee an AI will not harm humans, so we should be careful about what capabilities we give them. We do not controlneural networks well… AR Armin Ronacher — holds since 2023-02-17 Running coding agents without approval prompts is a normalization-of-deviance risk: the fact that it has not caused a disaster yet is itself the problem. Running coding agentswithout approval… SW Simon Willison — holds since 2025-12-31 Truly self-sovereign AI agents, running on compute no human owner can switch off, are inevitable rather than merely possible. Truly self-sovereign AIagents, running on… DB Dean W. Ball — holds since 2026-09-01 AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. AI agents werereasonable to assume a… ZM Zvi Mowshowitz — holds since 2026-08-29 AI labs tend to downplay safety risks in their public rhetoric even while acting responsibly in practice. AI labs tend todownplay safety risks… ZM Zvi Mowshowitz — holds since 2026-09-04 In an agent-written codebase the agent-written tests are equally untrustworthy; manually using the product is the only reliable measure of whether it works. In an agent-writtencodebase the… MZ Mario Zechner — holds since 2026-03-25 Confidence in safety, not capability, will increasingly set the pace of AI progress. Confidence in safety,not capability, will… SA Sam Altman — holds since 2026-08-18
same subject