Corrupt AI Model
How AML.T0076 Corrupt AI Model shows up in practice: the mapped risk classes in this atlas, the documented incidents that prove it's real, and the scenarios and controls to learn and defend against it.
Mapped risks
Risk classes in this atlas that map to AML.T0076 — click through for the full definition, attack surface and controls.
Real-world cases
2Documented incidents, disclosed vulnerabilities and research that illustrate AML.T0076 — latest first, each with sources.
Heretic automates 'abliteration' — removing an open model's safety refusals by orthogonalizing the refusal direction out of its weights, with an Optuna search that preserves capability — and has produced 4000+ uncensored models on Hugging Face.
Safety refusals in open models can be removed via a single-direction edit; '-abliterated' uncensored models then proliferated on public hubs.
Controls & guardrails that address this
3Guardrails across the risks mapped to AML.T0076, grouped by control function. Filter by control category below.
Knowing exactly where the model came from, checking it hasn't been swapped, and testing its behaviour before going live.
Regularly testing the AI against a set of known-good and known-bad examples, and re-testing whenever anything changes.
The organisational habits around the AI: assessing risks before launch, actively trying to break it, and having a plan for when something goes wrong.