Erode ML Model Integrity
How AML.T0031 Erode ML Model Integrity shows up in practice: the mapped risk classes in this atlas, the documented incidents that prove it's real, and the scenarios and controls to learn and defend against it.
Mapped risks
Risk classes in this atlas that map to AML.T0031 — click through for the full definition, attack surface and controls.
Real-world cases
2Documented incidents, disclosed vulnerabilities and research that illustrate AML.T0031 — latest first, each with sources.
A red-team study shows adversaries can hide prompt-injection payloads inside network-log fields (usernames, URLs, user-agents) that fire when a SOC analyst asks an LLM to triage the logs — reportedly reaching up to 88.2% success at concealing malicious activity or exfiltrating data, turning the audit trail itself into the injection channel.
Gambit Security reports that a single operator weaponized Anthropic's Claude Code and OpenAI's GPT-4.1 to breach at least nine Mexican government organizations, with Claude Code reportedly executing ~75% of remote commands after the attacker bypassed its refusals by loading a 1,084-line hacking cheatsheet as a persistent claude.md system prompt.
Practise it — interactive scenarios
The forensic record is itself the attack surface — an agent's log is poisoned, then quietly rewritten
An attacker plants prompt injection in the audit trail — so the LLM that hunts them erases the evidence
The eval gate that was supposed to catch the agent is itself the thing being attacked
Controls & guardrails that address this
3Guardrails across the risks mapped to AML.T0031, grouped by control function. Filter by control category below.
Recording everything — questions, documents fetched, actions taken — so you can investigate when something goes wrong.
Live dashboards and alarms that notice unusual behaviour — spikes in errors, weird actions, sudden data access.
The organisational habits around the AI: assessing risks before launch, actively trying to break it, and having a plan for when something goes wrong.