← Frameworks
AML.T0043.004

Craft Adversarial Data: Insert Backdoor Trigger

How AML.T0043.004 Craft Adversarial Data: Insert Backdoor Trigger shows up in practice: the mapped risk classes in this atlas, the documented incidents that prove it's real, and the scenarios and controls to learn and defend against it.

Real-world cases

6

Documented incidents, disclosed vulnerabilities and research that illustrate AML.T0043.004 — latest first, each with sources.

TeamPCP poisons the LiteLLM AI gateway on PyPI to harvest LLM API keys24 Mar 2026

As part of a multi-ecosystem supply-chain cascade (Trivy onward), TeamPCP used stolen PyPI publishing tokens to ship backdoored BerriAI LiteLLM versions whose auto-running .pth payload harvested cloud, SSH and Kubernetes secrets plus env vars holding OPENAI_API_KEY/ANTHROPIC_API_KEY — exfiltrating to a typosquatted C2; AI-talent firm Mercor was a downstream victim, with Lapsus$ claiming ~4TB stolen.

ClawHavoc — mass poisoning of OpenClaw's ClawHub agent-skill marketplace01 Feb 2026

Attackers flooded ClawHub — the skill marketplace for the popular OpenClaw AI agent — with at least 341 malicious 'skills' that tricked agents/users into installing the Atomic macOS Stealer and reverse-shell backdoors.

A small number of samples can poison LLMs of any size (~250-document backdoor)08 Oct 2025

Anthropic, the UK AI Security Institute and the Alan Turing Institute report that a near-constant number of poisoned documents (~250 in their experiments) reliably installs a backdoor in models from 600M to 13B parameters — suggesting poisoning cost may be a roughly fixed absolute count rather than a percentage of training data. The authors stress the demonstrated backdoor is narrow (a denial-of-service trigger) and likely not a frontier-model risk on its own.

Malice in Agentland — backdooring agents through the supply chain (Boisvert et al.)03 Oct 2025 (rev. 2026)

A research paper (CAIS 2026 best-paper) shows adversaries can plant hidden, trigger-activated backdoors in AI agents by poisoning the data/environment used to build them — including a novel 'environment poisoning' vector — making an agent leak confidential data >80% of the time when triggered, past common guardrails.

Sleeper Agents (Hubinger et al., Anthropic)10 Jan 2024

Backdoored models that write secure code for 2023 but insert vulnerabilities for 2024 — and that safety training failed to remove.

PoisonGPT (Mithril Security)09 Jul 2023

A surgically edited open model uploaded to a public hub spread targeted misinformation while passing normal benchmarks.

Browse all real-world cases →

Controls & guardrails that address this

3

Guardrails across the risks mapped to AML.T0043.004, grouped by control function. Filter by control category below.

Control category
Preventive · 1
Weight provenance, hashing & pre-deploy evalsinteractive

Knowing exactly where the model came from, checking it hasn't been swapped, and testing its behaviour before going live.

Open the Control Library →

AI RiskAtlas is an educational model of how GenAI & agentic systems work and fail. Architectures and payloads are illustrative and simplified for learning — not operational guidance. Real-world cases are summarised from public reporting.

Sources & further reading →·Built by Shi Yuan ↗