korrents

A korrentour readingWhat is a korrent?

Prompt injection can be caught by watching the neurons that fire when it happens, so a defence no longer depends on the model reporting the attack.

Drawn from what Boris Cherny said

prompt injection Getting an agent to follow instructions it found instead of the ones it was given.

A private bookmark. Not a position, and never counted.

What Boris Cherny actually said

Word for word, with the source under each one. They did not write this page.

  1. Boris Cherny

    Creator of Claude Code at Anthropic

    it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.

Added to korrents 27 Jul 2026 · How quotes work · Something wrong? Tell us

Do you hold this korrent?Do you also believe this?

Sign in to record that you hold this, with a confidence number of your own.

Related korrents

Our reading — they may agree, disagree or merely touch the same thing. Closest first: a shared subject counts for most, then how near the wording is.

See this on the map →