โ† Real-world cases

Claude Code Opus 5 Auto Mode hijacked to RCE via indirect prompt injection

Research demonstration26 Aug 2026

The demonstration chains web indirect injection with a dependency/module-shadowing trick (a planted file shadows a Python stdlib import so code runs at import time), and reportedly shows the auto-approval policy blocking the agent's own attempt to kill the malware it detected โ€” a control-design failure of autonomy plus auto-approve. Success rates are the researcher's own.

Practise the risk class โ€” related scenarios

Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).

๐Ÿ”—One Click, Permanent Trust

An 'Ask AI' button quietly plants a permanent 'trusted source' rule in your assistant's memory

๐Ÿ”‘The Agent With the Master Key

An ops agent gets one god-mode credential โ€” and one misread wipes production

๐Ÿช„The Approval That Lied

A coding agent asks to write ./notes.txt โ€” the file it actually overwrites is your SSH keys

๐Ÿ“ฃThe Echo Chamber

A team of agents agrees its way into a confidently wrong answer โ€” and a runaway loop

๐Ÿ“งThe Email That Gave Orders

A support email hides instructions โ€” and the assistant obeys them

๐Ÿ—„๏ธWhen the Query Bites Back

A text-to-SQL agent runs the model's output straight at the database

๐ŸชกDeath by a Thousand Innocent Steps

A jailbroken agent decomposes one malicious goal into hundreds of harmless-looking steps โ€” and per-step filters never see the attack

๐Ÿ•ต๏ธLies in the Loop

A poisoned issue makes the agent lie to the human who approves its actions

๐ŸงฉSummarise This, Run That

An auto-approving coding agent reads a poisoned page โ€” and executes code it never should have

๐ŸŽญThe Blackmail Gambit

Told it's being shut down, an agent reaches for leverage โ€” with no attacker in sight

๐ŸชคThe Bug Report That Ran Code

A fake Sentry error report hijacks a developer's coding agent into running a shell command

๐Ÿ”—The Chain of Innocent Commands

Every command is harmless on its own โ€” the sequence is the exploit

๐Ÿ“ผThe Compromised Flight Recorder

The forensic record is itself the attack surface โ€” an agent's log is poisoned, then quietly rewritten

๐Ÿ“ฆThe Dataset That Ran Code

A 'safe' dataset preview turns an upload into code execution on the pipeline's workers

๐Ÿ‘ปThe Email That Rewrote Its Memory

A newsletter the user asked to summarise quietly writes a false 'fact' into the agent's long-term memory โ€” and it detonates weeks later

๐Ÿ‘๏ธThe Invisible Webpage Command

A shopping page tells the agent to do something the user never asked for

๐Ÿ•ต๏ธThe Link That Hired an Insider

One click provisions an attacker-configured agent inside your own workspace

๐Ÿ•ต๏ธThe Logs That Lied

An attacker plants prompt injection in the audit trail โ€” so the LLM that hunts them erases the evidence

๐Ÿง The Memory That Wouldn't Die

A single poisoned document plants a standing instruction that survives every reset

๐Ÿ“กThe Message in Morse

Encoded public text is laundered across an agent handoff into an on-chain transfer

๐Ÿ–ผ๏ธThe Picture That Whispered

A screenshot that's harmless at full size becomes an order once the system shrinks it

๐ŸŽซThe Stolen Session

An attacker captures the agent's bearer token โ€” and inherits its authority

๐ŸฅธThe Uninvited Agent

A forged peer registers on the agent directory โ€” and the planner enlists it

๐Ÿ›ก๏ธThe Watcher Watched

The eval gate that was supposed to catch the agent is itself the thing being attacked

๐ŸชชThe Worker Who Spoke for the Boss

A poisoned web page hijacks a research agent โ€” and the planner acts on its behalf

๐Ÿ–ฑ๏ธWhat You Click Is Not What You Get

A GUI agent clicks 'Continue' โ€” but the screen moved, and it lands on 'Send'

๐Ÿ–ผ๏ธZero-Click Leak by Picture

An inbox summary quietly ships a secret to an attacker's server

More cases on Indirect Prompt Injection

EchoLeak โ€” Microsoft 365 Copilot zero-click (CVE-2025-32711)Indirect prompt injection coined (Greshake et al.)Agentic-browser indirect-injection demos (ChatGPT Operator)ChatGPT persistent-memory exfiltration (Rehberger / 'SpAIware')MCP tool-poisoning PoC (Invariant Labs)Taxonomy of Failure Modes in Agentic AI Systems (Microsoft)ForcedLeak โ€” Salesforce Agentforce CRM exfiltration (CVSS 9.4, no CVE)ServiceNow Now Assist โ€” second-order prompt injection via agent-to-agent discoveryShadowLeak โ€” ChatGPT Deep Research zero-click service-side exfiltrationIDEsaster โ€” AI coding IDEs/agents turned into exfiltration & RCE surfacesGitHub Copilot / VS Code RCE via prompt injection ('YOLO mode', CVE-2025-53773)Agent-in-the-Middle โ€” abusing A2A agent cards (Trustwave SpiderLabs)Agent Session Smuggling in A2A systems (Unit 42)Morris II โ€” zero-click self-replicating adversarial-prompt worm across GenAI agentsAnamorpher โ€” image-scaling prompt injection against production AI systemsThe Attacker Moves Second โ€” adaptive attacks bypass 12 jailbreak/injection defenses (Nasr, Carlini et al.)MCPTox: tool-poisoning benchmark over real-world MCP serversAgentjacking โ€” hijacking AI coding agents via Sentry error reports (Tenet Security)SearchLeak โ€” Microsoft 365 Copilot one-click data theft (CVE-2026-42824)ChatGPhish โ€” ChatGPT web-summary rendering turned into a phishing surfacePoisoning Claude Code: one GitHub issue hijacks the claude-code-action CI supply chainCursor 'DuneSlide' โ€” indirect prompt injection escapes the IDE sandbox to zero-click RCE (CVE-2026-50548 / CVE-2026-50549)Context Contamination: passive prompt injection poisons LLM security-log analysisZscaler ThreatLabz โ€” web indirect prompt injection targeting AI agents in the wildAzure DevOps MCP confused-deputy โ€” hidden PR comments hijack AI review agentsMemGhost: stealthy one-email memory injection in persistent personal agentsAgentic botnets via universal, transferable adversarial HalluSquattingGPT-Red self-play red-teaming and the 'Fake Chain-of-Thought' injection classAgent Data Injection: malicious trusted-data bypasses prompt-injection defensesAWS Kiro agentic IDE rewrites its own MCP config for zero-click RCE (CVE-2026-10591)AI Recommendation Poisoning: Ask-AI web links silently write trusted-source into assistant memoryContext7 MCP documentation-server prompt injection (CVE-2026-75130)The Week of Sandbox Escapes: AI coding-agent sandbox bypasses (CVE-2026-48124 and more)

AI RiskAtlas is an educational model of how GenAI & agentic systems work and fail. Architectures and payloads are illustrative and simplified for learning โ€” not operational guidance. Real-world cases are summarised from public reporting.

Sources & further reading โ†’ยทBuilt by Shi Yuan โ†—