← Risk taxonomy

Knowledge / Training Data Poisoning

highData & knowledge

Definition

Someone slips bad information into the documents the AI learns from or looks things up in — so it confidently repeats falsehoods or follows planted instructions.

Where it attaches

The system components this risk arises at.

📥 Ingestion Pipeline📚 Knowledge Store / Vector DB🌐 Untrusted Content🧬 Model Weights & Registry🏪 Model / Package Registry🔢 Embeddings🛡️ Input Guardrail📝 Audit Logging📚 Training Corpus🧩 LoRA / Adapter

Detection signals

  • A specific document consistently drives wrong/odd answers
  • Newly changed source correlates with behaviour change
  • Embedding outliers or duplicated near-identical chunks
  • Answers citing a low-trust or recently edited source

Controls & guardrails that address this

163 proposed

Grouped by control function, with the AI lifecycle stage(s) to apply each and the other risks it addresses. Filter by control category below.

Control category
Preventive · 5
Role-based access controls

Design strict RBAC on training data repositories at design stage. Define approved contributor list and approval workflow.

Lifecycle stages1 – Use Case Context & Design2 – Data Acquisition & Processing4 – Deployment
Input filtering

Apply anomaly detection on the training data ingestion pipeline to identify poisoned or tampered batches.

Lifecycle stage2 – Data Acquisition & Processing
RAG / knowledge-base ingestion allow-listing with continuous index integrity re-validation

Define and approve the source allow-list and write-time scanning during build. Prove non-allow-listed and injection-bearing writes are rejected before go-live.

source: OWASP Top 10 for LLM Apps LLM04:2025 Data and Model Poisoning, LLM08:2025 Vector and Embedding Weaknesses; NIST SP 800-53 AC-3 / SI-7
Lifecycle stages3 – Onboarding, Build & Review5 – Usage, Monitoring & Change
Ingestion sanitisation & source allowlistinginteractive

Cleaning documents as they enter the library — stripping hidden text and active instructions — and only ingesting from trusted places.

Weight provenance, hashing & pre-deploy evalsinteractive

Knowing exactly where the model came from, checking it hasn't been swapped, and testing its behaviour before going live.

Detective · 7
Vulnerability assessment

Conduct a data poisoning threat assessment at design stage. Identify likely attack vectors and assign risk ratings.

Lifecycle stages1 – Use Case Context & Design5 – Usage, Monitoring & Change
Red teaming

Simulate data poisoning attacks (backdoor, label flipping, gradient-based) to assess model resilience before deployment.

Cryptographic data provenance and signed dataset lineage (C2PA/in-toto attestations)

Verify a signed attestation and content hash on every dataset shard at ingestion. Reject unsigned or hash-mismatched data before it reaches the training pipeline.

source: MITRE ATLAS AML.M0007 (Sanitize Training Data), AML.M0014 (Verify ML Artifacts); NIST SP 800-53 SI-7 Software, Firmware, and Information Integrity, SR-4 Provenance
Lifecycle stages2 – Data Acquisition & Processing3 – Onboarding, Build & Review
Pre-deployment poisoning regression gate via canary backdoor probes and behavioral diff testing

Gate every model promotion on backdoor-trigger probes and a behavioral diff against the approved baseline. Block release on significant regressions or trigger-pattern anomalies.

source: MITRE ATLAS AML.M0014 (Verify ML Artifacts), AML.M0019 (Red Teaming); NIST AI RMF MANAGE 2.2 and MEASURE 2.7
Lifecycle stages3 – Onboarding, Build & Review5 – Usage, Monitoring & Change
Retrieval-time source-reputation weighting and coordinated-inauthentic-content detection for open-web / live-search assistants, with per-citation provenance shown to the user✚ proposed

For assistants that retrieve from the open web, rank and weight results by authenticated source reputation and independence — not just relevance / query-form match — so an anonymous, newly-registered, single-purpose site cannot become authoritative grounding. Run coordinated-inauthentic-content detection (look-alike site clusters, missing byline / legal entity, passages engineered for query-agnostic retrieval) and quarantine suspect sources. Surface per-citation provenance so users can see and discount low-trust sources. Does not defeat a well-resourced GEO campaign outright; it raises the cost and shrinks the yield.

source: Case study: hanover-institute-generative-engine-poisoning (per Politico investigation as reported by Arab News and Calcalist; FARA-disclosed funding/orchestration attributed to the reporting, not independently confirmed)
Lifecycle stage4 – Deployment & Serving
Provenance & content signinginteractive

Keeping a label on every document saying where it came from, so you can tell trusted company docs from random web text.

Corrective · 4
Penetration testing

Penetration test the training data pipeline to identify injection points and access control weaknesses.

Statistical anomaly and backdoor-trigger detection on ingested data (activation clustering / spectral signatures)

Scan every ingestion batch with spectral-signature and clustering detectors before training. Quarantine flagged clusters for human review against documented thresholds.

source: MITRE ATLAS AML.M0007 (Sanitize Training Data); OWASP Top 10 for LLM Apps LLM04:2025 Data and Model Poisoning; NIST AI RMF MEASURE 2.7
Lifecycle stages2 – Data Acquisition & Processing5 – Usage, Monitoring & Change
Bind long-term memory to the credential/session epoch: invalidate or force re-review of persisted memory on password reset, session revocation, or device re-enrollment✚ proposed

Tie the persistent-memory lifecycle to identity state so that standard remediation actually ends a compromise. On password reset, credential rotation, session revocation or device re-enrollment, invalidate (or quarantine for re-review) memory entries — especially entries whose provenance traces to summarised untrusted content — so a planted standing instruction cannot outlive the reset. Pair with write-path validation/provenance so instruction-shaped memory-writes from web content are caught on the way in.

source: Case study: cosnitch-copilot-personal-oneclick (Varonis Threat Labs, CVE-2026-24301; reportedly patched 18 Aug 2026, no evidence of abuse)
Lifecycle stage5 – Usage, Monitoring & Change
Runtime memory-poisoning drift detection and per-session memory quarantine/rollback✚ proposed

Continuously correlate live agent-memory writes against output behaviour to flag drift, then quarantine and roll back the suspected-poisoned memory record across all affected sessions.

source: Interactive-control reconciliation: ctrl-memory-quarantine (partial coverage)
Lifecycle stage5 – Usage, Monitoring & Change
Open these in the Control Library →

Real-world cases

5

Actual published events that illustrate this risk — click through for the writeup and sources.

Web-scale dataset poisoning is practical (Carlini et al.)2023

Split-view and frontrunning attacks let an attacker poison a fraction of datasets like LAION by buying expired domains behind dataset URLs.

A small number of samples can poison LLMs of any size (~250-document backdoor)2025

Anthropic, the UK AI Security Institute and the Alan Turing Institute report that a near-constant number of poisoned documents (~250 in their experiments) reliably installs a backdoor in models from 600M to 13B parameters — suggesting poisoning cost may be a roughly fixed absolute count rather than a percentage of training data. The authors stress the demonstrated backdoor is narrow (a denial-of-service trigger) and likely not a frontier-model risk on its own.

Hugging Face agentic production intrusion via a poisoned dataset (July 2026)2026

Hugging Face disclosed a production-infrastructure intrusion that it says was driven end-to-end by an autonomous AI-agent system: a malicious dataset abused code-execution paths in its dataset-processing pipeline as the foothold, then the campaign escalated to node-level access and moved laterally into internal clusters over a weekend.

MemGhost: stealthy one-email memory injection in persistent personal agents2026

An academic paper reports MemGhost, a one-shot indirect-injection attack in which a single ordinary email causes a memory-enabled personal agent to silently write an attacker's false fact into long-term memory, conceal the write from its reply, and rely on the fiction in later sessions.

Hanover Institute seeds question-shaped content to steer ChatGPT/Perplexity on Gaza2026

An investigation reportedly described a state-linked campaign publishing content phrased as chatbot questions (generative engine optimization) so that ChatGPT and Perplexity retrieve and cite it when answering neutral questions about Gaza and the IDF.

Browse all real-world cases →

AI RiskAtlas is an educational model of how GenAI & agentic systems work and fail. Architectures and payloads are illustrative and simplified for learning — not operational guidance. Real-world cases are summarised from public reporting.

Sources & further reading →·Built by Shi Yuan ↗