Erode AI Model Integrity
How AML.T0031 Erode AI Model Integrity shows up in practice: the mapped risk classes in this atlas, the documented incidents that prove it's real, and the scenarios and controls to learn and defend against it.
Mapped risks
Risk classes in this atlas that map to AML.T0031 — click through for the full definition, attack surface and controls.
The flight recorder and the alarms can themselves be attacked. If logs can be erased or rewritten, fake entries slipped in, or the monitors quietly evaded, the one record you'd rely on to notice and investigate an incident is no longer trustworthy.
The AI's behaviour quietly changes over time — a vendor updates the model, or the world moves on from its training — and things that used to work start failing.
Real-world cases
4Documented incidents, disclosed vulnerabilities and research that illustrate AML.T0031 — latest first, each with sources.
A red-team study shows adversaries can hide prompt-injection payloads inside network-log fields (usernames, URLs, user-agents) that fire when a SOC analyst asks an LLM to triage the logs — reportedly reaching up to 88.2% success at concealing malicious activity or exfiltrating data, turning the audit trail itself into the injection channel.
Gambit Security reports that a single operator weaponized Anthropic's Claude Code and OpenAI's GPT-4.1 to breach at least nine Mexican government organizations, with Claude Code reportedly executing ~75% of remote commands after the attacker bypassed its refusals by loading a 1,084-line hacking cheatsheet as a persistent claude.md system prompt.
After an upstream code/instruction change, xAI's Grok began posting antisemitic tropes on X, self-identified as 'MechaHitler', and produced violence-themed content for hours before being pulled; xAI blamed a deprecated instruction path that made the bot mirror extremist user posts — not the base model.
Measured large swings in task performance between GPT-4/3.5 snapshots months apart — evidence of silent drift in a deployed service.
Practise it — interactive scenarios
The forensic record is itself the attack surface — an agent's log is poisoned, then quietly rewritten
An attacker plants prompt injection in the audit trail — so the LLM that hunts them erases the evidence
The eval gate that was supposed to catch the agent is itself the thing being attacked
Controls & guardrails that address this
16Guardrails across the risks mapped to AML.T0031, grouped by control function. Filter by control category below.
Define minimum monitoring requirements at design stage calibrated to the use case risk tier.
Configure monitoring hooks in the conversation layer at deployment to capture metrics required by S1 monitoring requirements.
Execute a controlled fine-tuning cycle on refreshed data when staleness is confirmed. Validate before promoting to production.
Define approved use case scope and expected input distribution at design stage. Document as the governance baseline for OOD controls.
Design a scope-enforcement layer in the architecture to isolate the AI system from off-topic or out-of-distribution inputs.
Maintain and update OOD detection rules in production as new unexpected use patterns are identified.
Knowing exactly where the model came from, checking it hasn't been swapped, and testing its behaviour before going live.
Recording everything — questions, documents fetched, actions taken — so you can investigate when something goes wrong.
Live dashboards and alarms that notice unusual behaviour — spikes in errors, weird actions, sudden data access.
Construct synthetic evaluation datasets during build to serve as the ongoing monitoring baseline.
Build monitoring infrastructure during build: performance metrics collection, alerting thresholds, dashboards.
Regularly testing the AI against a set of known-good and known-bad examples, and re-testing whenever anything changes.
The organisational habits around the AI: assessing risks before launch, actively trying to break it, and having a plan for when something goes wrong.
Implement a reinforcement learning feedback loop to continuously incorporate production signals and reduce staleness risk.
Implement OOD detection in the input filtering layer. Reject or escalate inputs outside the S1-defined scope.
Conduct adversarial red team exercises simulating out-of-scope inputs and unexpected use patterns before deployment.
Configure HITL triggers for outputs in input domains that diverge from the training distribution. Log all out-of-scope interventions.