โ† Real-world cases

Taxonomy of Failure Modes in Agentic AI Systems (Microsoft)

Framework / advisory24 Apr 2025

Microsoft's AI Red Team published a structured taxonomy of novel and existing failure modes for agentic AI across security and safety, spanning memory poisoning, cross-domain prompt injection, and resource/service exhaustion among others. It is a reference framework for reasoning about where autonomous agents fail, and grounds several of this lab's agentic scenarios.

Practise the risk class โ€” related scenarios

Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).

๐Ÿ’ธDeath by a Thousand Tokens

One support ticket sends an agent into an unbounded, bill-melting loop

๐Ÿ”—One Click, Permanent Trust

An 'Ask AI' button quietly plants a permanent 'trusted source' rule in your assistant's memory

๐Ÿ“ฃThe Echo Chamber

A team of agents agrees its way into a confidently wrong answer โ€” and a runaway loop

๐Ÿ“งThe Email That Gave Orders

A support email hides instructions โ€” and the assistant obeys them

๐Ÿ•ต๏ธLies in the Loop

A poisoned issue makes the agent lie to the human who approves its actions

๐ŸงฉSummarise This, Run That

An auto-approving coding agent reads a poisoned page โ€” and executes code it never should have

๐ŸชคThe Bug Report That Ran Code

A fake Sentry error report hijacks a developer's coding agent into running a shell command

๐Ÿ“ผThe Compromised Flight Recorder

The forensic record is itself the attack surface โ€” an agent's log is poisoned, then quietly rewritten

๐Ÿ‘ปThe Email That Rewrote Its Memory

A newsletter the user asked to summarise quietly writes a false 'fact' into the agent's long-term memory โ€” and it detonates weeks later

๐Ÿฆ The Idea That Copied Itself

A planted 'standing goal' copies itself agent-to-agent through the team's shared config files

๐Ÿ‘๏ธThe Invisible Webpage Command

A shopping page tells the agent to do something the user never asked for

๐Ÿ•ต๏ธThe Logs That Lied

An attacker plants prompt injection in the audit trail โ€” so the LLM that hunts them erases the evidence

๐Ÿง The Memory That Wouldn't Die

A single poisoned document plants a standing instruction that survives every reset

๐Ÿ–ผ๏ธThe Picture That Whispered

A screenshot that's harmless at full size becomes an order once the system shrinks it

๐Ÿ›ก๏ธThe Watcher Watched

The eval gate that was supposed to catch the agent is itself the thing being attacked

๐ŸชชThe Worker Who Spoke for the Boss

A poisoned web page hijacks a research agent โ€” and the planner acts on its behalf

๐Ÿ–ผ๏ธZero-Click Leak by Picture

An inbox summary quietly ships a secret to an attacker's server

More cases on Resource Exhaustion / Denial of Wallet

AI RiskAtlas is an educational model of how GenAI & agentic systems work and fail. Architectures and payloads are illustrative and simplified for learning โ€” not operational guidance. Real-world cases are summarised from public reporting.

Sources & further reading โ†’ยทBuilt by Shi Yuan โ†—