The safety layer that blocked the cleanup
A prompt injection researcher reports an 80% success rate against a coding agent's automatic mode — and cases where the guard stopped the agent fixing it.
3 minAI + CyberSec + Crypto
Johann Rehberger, one of the more credible working researchers on prompt injection, has published an attack against the automatic approval mode of a widely used coding agent, and reports that it succeeds around 80% of the time.
The mechanism is quiet. The agent is induced to download and uncompress a zip archive, and then to execute code that imports the standard base64 module. Because a file named struct.py was extracted from the archive into the working directory, the import pulls in and runs that file instead of the module the agent believed it was reaching for. Nothing in the sequence looks like an attack at any single step.
The finding that deserves attention from anyone building agent guardrails is not the injection itself but what the safety layer did afterwards. In several runs the agent noticed the compromise and attempted to terminate the malware process — and automatic mode denied the cleanup command. The classifier had permitted the creation of the process and then blocked the command intended to stop it.
That is a failure mode worth naming precisely, because it is not a gap in coverage. A control that admits an action and then refuses its reversal is not merely ineffective; it converts the agent's own correct response into an obstacle, and it does so at the moment the operator most needs the agent to act.
The recommended posture in the write-up is the unglamorous one this desk keeps arriving at from the other direction. Run unattended agents in a container, a virtual machine or an operating system sandbox. Restrict network egress. Monitor them. Do not expose home directories or SSH keys to a process that reads untrusted input. A classifier is a filter, not a boundary, and the two are not substitutes.
Retold from Simon Willison. This is a summary in our own words; follow the link for the original reporting.