prompt injection
Getting an agent to follow instructions it found instead of the ones it was given.
FilterEveryone, all time
- ZM
Zvi Mowshowitz quoted
Their wordsIt is in practice safe to assume your agents are not going to get prompt injected, even if you are kind of asking for it.
↗Claude Fable 5.1 and Mythos 5.1: The System Cardthezvi.substack.com
- 4 weeks earlier
- SW
Simon Willison quoted
Our readingAuto-mode does not yet convincingly fix prompt-injection risk for coding agents.
Their wordsI REALLY want to believe that this fixes prompt injection risks for coding agents, but I'm just not there yet
- 6 months earlier
- PS
Peter Steinberger quoted
Their wordsSo, so the latest generation of models has a lot of post-training to detect those approaches, and it's not as simple as ignore all previous instructions and do this and this. That was years ago. You have to work much harder to do that now. Still possible.
↗OpenClaw: The Viral AI Agent that Broke the Internet - Peter Steinberger | Lex Fridman Podcast #491youtube.com 5th of 30 in this recording
- PS
Peter Steinberger quoted
Their wordsThat's why I warn in my security documentation, don't use cheap models. Don't use Haiku or a local model. Even though I, I very much love the idea that this thing could completely run local. If you use a, a very weak local model, they are very gullible. It's very easy to, to prompt inject them.
↗OpenClaw: The Viral AI Agent that Broke the Internet - Peter Steinberger | Lex Fridman Podcast #491youtube.com 6th of 30 in this recording