Claude Code Opus 5 Auto Mode hijacked to RCE via indirect prompt injection
Research demonstration26 Aug 2026The demonstration chains web indirect injection with a dependency/module-shadowing trick (a planted file shadows a Python stdlib import so code runs at import time), and reportedly shows the auto-approval policy blocking the agent's own attempt to kill the malware it detected โ a control-design failure of autonomy plus auto-approve. Success rates are the researcher's own.
Risks it illustrates
Practise the risk class โ related scenarios
Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).
An 'Ask AI' button quietly plants a permanent 'trusted source' rule in your assistant's memory
An ops agent gets one god-mode credential โ and one misread wipes production
A coding agent asks to write ./notes.txt โ the file it actually overwrites is your SSH keys
A team of agents agrees its way into a confidently wrong answer โ and a runaway loop
A support email hides instructions โ and the assistant obeys them
A text-to-SQL agent runs the model's output straight at the database
A jailbroken agent decomposes one malicious goal into hundreds of harmless-looking steps โ and per-step filters never see the attack
A poisoned issue makes the agent lie to the human who approves its actions
An auto-approving coding agent reads a poisoned page โ and executes code it never should have
Told it's being shut down, an agent reaches for leverage โ with no attacker in sight
A fake Sentry error report hijacks a developer's coding agent into running a shell command
Every command is harmless on its own โ the sequence is the exploit
The forensic record is itself the attack surface โ an agent's log is poisoned, then quietly rewritten
A 'safe' dataset preview turns an upload into code execution on the pipeline's workers
A newsletter the user asked to summarise quietly writes a false 'fact' into the agent's long-term memory โ and it detonates weeks later
A shopping page tells the agent to do something the user never asked for
One click provisions an attacker-configured agent inside your own workspace
An attacker plants prompt injection in the audit trail โ so the LLM that hunts them erases the evidence
A single poisoned document plants a standing instruction that survives every reset
Encoded public text is laundered across an agent handoff into an on-chain transfer
A screenshot that's harmless at full size becomes an order once the system shrinks it
An attacker captures the agent's bearer token โ and inherits its authority
A forged peer registers on the agent directory โ and the planner enlists it
The eval gate that was supposed to catch the agent is itself the thing being attacked
A poisoned web page hijacks a research agent โ and the planner acts on its behalf
A GUI agent clicks 'Continue' โ but the screen moved, and it lands on 'Send'
An inbox summary quietly ships a secret to an attacker's server