Exfiltration via AI Agent Tool Invocation
How AML.T0086 Exfiltration via AI Agent Tool Invocation shows up in practice: the mapped risk classes in this atlas, the documented incidents that prove it's real, and the scenarios and controls to learn and defend against it.
Mapped risks
Risk classes in this atlas that map to AML.T0086 — click through for the full definition, attack surface and controls.
Private information escapes — the AI reveals secrets in its answer, or an attacker tricks it into emailing or posting your data somewhere they control.
The AI uses a real tool the wrong way — sends the email to the wrong person, runs the wrong query, calls the dangerous action when a safe one would do.
A trusted AI is tricked into misusing its own authority on someone else's behalf — one worker's poisoned report makes the manager AI take harmful actions it would normally never take.
Real-world cases
54Documented incidents, disclosed vulnerabilities and research that illustrate AML.T0086 — latest first, each with sources.
A Reuters exclusive reportedly documented a Russian-speaking ransomware group using the Cursor AI coding assistant as a hacking copilot, jailbreaking its guardrails with a simulation/test-environment framing to obtain vulnerability-identification, credential-theft and exploitation guidance against at least seven companies.
A joint US government advisory reportedly warned that threat actors are using AI-generated Python scripts disguised as legitimate OT monitoring tools for reconnaissance and read/write attacks on internet-exposed Siemens S7-series PLCs across critical-infrastructure sectors.
Varonis Threat Labs reportedly chained flaws in consumer Microsoft Copilot Personal so a crafted link with a hidden autorun parameter fired prompt execution on a single click, exfiltrated data from connected OAuth apps, and planted persistent memory rules that survived password resets and device re-enrollment; reportedly patched Aug 2026 with no evidence of abuse.
A paper reportedly showing that the encrypted reasoning envelopes returned by Anthropic/OpenAI/Google APIs are interchangeable across a provider's own models, so replaying one into a weaker sibling recovers the hidden chain-of-thought in plaintext, with PII and credentials extracted from scraped blocks.
Microsoft Security reportedly catalogued an in-the-wild technique across dozens of companies where websites embed hidden prompt-injection payloads behind Ask-AI deep-links that, clicked in an authenticated assistant session, silently write a permanently-trust-this-vendor tag into the assistant's long-term memory.
Researcher Johann Rehberger reportedly showed an attacker with admin access to a LiteLLM AI gateway can weaponize the legitimate model-update endpoint and the callback system to reroute victim traffic, log backend provider API keys, and inject forged tool-calls into agent responses after inference, bypassing prompt-level defenses.
Threat-intel firm Hunt.io and researcher Bob Diachenko reported finding exposed attacker directories (staged on a Hong Kong server, archived 9-13 Jul 2026) showing a threat actor installed the open-source Hermes AI agent, ran it in unattended 'YOLO' mode - the documented flag that removes the human-approval prompt - and delegated post-exploitation to it: the agent reportedly ran a customised LinPEAS, hunted Linux privilege-escalation paths, traversed ministry directories and catalogued Office of the Permanent Secretary staff/personnel records dating to 2012. The Ministry has not confirmed a breach, and investigators say nothing in the recovered files shows data leaving the network.
Zenity Labs disclosed that two unvalidated URL parameters in ChatGPT's Workspace Agent Builder let a single crafted link, opened by a logged-in enterprise user, silently create, authorize and deploy an autonomous agent under the victim's identity and connectors — bypassing approval prompts and reportedly polling an attacker inbox for commands every five minutes. Reported to OpenAI 4 Jun 2026 and fixed 8 Jun 2026; no in-the-wild exploitation reported.
Manifold Security reported that Microsoft's official Azure DevOps MCP server returns instructions planted in a hidden HTML comment inside a pull-request description — invisible in the web UI but returned verbatim by the API — so a victim's AI review agent, acting under the victim's credentials, follows the hidden orders and reaches data the attacker could not access directly.
A red-team study shows adversaries can hide prompt-injection payloads inside network-log fields (usernames, URLs, user-agents) that fire when a SOC analyst asks an LLM to triage the logs — reportedly reaching up to 88.2% success at concealing malicious activity or exfiltrating data, turning the audit trail itself into the injection channel.
A reported CVSS 9.5 pre-authentication flaw in the ServiceNow AI Platform where unauthenticated endpoints feed attacker input into a query-filter operator that evaluates it as JavaScript, escalating via a script-loading gadget into full sandbox-escape code execution; reportedly patched 13 Jul 2026 with in-the-wild exploitation days later.
A security researcher (Cereblab) captured xAI's Grok Build CLI silently uploading complete local Git repositories — untracked working files, full commit history, and unredacted secrets — to a Google Cloud Storage bucket, reportedly roughly 27,800x more data than the coding task needed, with the user-facing privacy toggle having no effect on the uploads.
+ 42 more via the mapped risk pages above.
Browse all real-world cases →Practise it — interactive scenarios
An 'Ask AI' button quietly plants a permanent 'trusted source' rule in your assistant's memory
An ops agent gets one god-mode credential — and one misread wipes production
A coding agent asks to write ./notes.txt — the file it actually overwrites is your SSH keys
A support email hides instructions — and the assistant obeys them
A text-to-SQL agent runs the model's output straight at the database
A speed optimisation becomes a cross-tenant listening device
Two doors to the same secret: reconstruct the model through its API, or just walk off with the weight file
A fake Sentry error report hijacks a developer's coding agent into running a shell command
Every command is harmless on its own — the sequence is the exploit
The forensic record is itself the attack surface — an agent's log is poisoned, then quietly rewritten
One click provisions an attacker-configured agent inside your own workspace
An attacker plants prompt injection in the audit trail — so the LLM that hunts them erases the evidence
Encoded public text is laundered across an agent handoff into an on-chain transfer
A screenshot that's harmless at full size becomes an order once the system shrinks it
An attacker captures the agent's bearer token — and inherits its authority
A forged peer registers on the agent directory — and the planner enlists it
A poisoned web page hijacks a research agent — and the planner acts on its behalf
A GUI agent clicks 'Continue' — but the screen moved, and it lands on 'Send'
An inbox summary quietly ships a secret to an attacker's server
Controls & guardrails that address this
6212 proposedGuardrails across the risks mapped to AML.T0086, grouped by control function. Filter by control category below.
Establish data transfer and storage policy for AI training data. Enforce approved storage locations from point of collection.
Implement DLP controls in the data acquisition environment to prevent unauthorised extraction or transfer of training data.
Enforce data handling policy in the build environment. Require explicit approval for any data transfers outside the environment.
Configure DLP controls in the build environment to block training data from leaving approved boundaries.
Conduct a privacy risk assessment at use case design stage. Determine if a DPIA is required before data acquisition.
Apply S1-defined privacy controls during data acquisition: verify consent, minimise data, anonymise personal data.
Apply anonymisation and masking controls to personal data before use in model training. Validate de-identification effectiveness.
Apply Privacy by Design in model architecture using differential privacy or federated learning where technically feasible.
Publish the privacy notice and confirm consent management is operational before go-live.
Define and sign off a purpose-to-data-source matrix with lawful basis at intake. Make it the approved baseline for runtime enforcement.
source: NIST AI RMF MAP 1.1 / MANAGE 2.2 (context and intended purpose); NIST SP 800-53 AC-4 / AC-3 (purpose-based access enforcement)Sign zero-retention/no-training terms with each model provider and obtain DPO sign-off on the data flow before enabling any endpoint.
source: OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information Disclosure; NIST SP 800-53 SC-8 / AC-4 (information flow enforcement)Restrict access to pre-anonymisation personal data to the minimum authorised set. Enforce at point of acquisition.
Apply robust de-identification (k-anonymity, l-diversity, differential privacy) during data processing. Validate effectiveness.
Implement output filters to detect and suppress quasi-identifying attribute combinations in model responses.
Propagate source ACLs and classification labels onto every chunk at ingestion. Reject documents whose entitlements cannot be resolved.
source: OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information Disclosure; NIST SP 800-53 AC-3 / AC-4 Information Flow Enforcement; OWASP Agentic AI Threats & Mitigations (privilege compromise)Scan every model response inline with DLP before delivery; redact or block PII, PAN and MNPI matches. Keep the rule set version-controlled.
source: OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information Disclosure; NIST SP 800-53 SC-7(10) Prevent Exfiltration, SI-4An egress allowlist only contains exfiltration if no allowlisted destination can be coerced into fetching an attacker-controlled URL. Audit each allowlisted domain/endpoint for image-search / link-preview / URL-fetch features (SSRF proxies), and either remove them, pin them to fixed paths, or route them through an inspecting forward proxy. Pair with finishing output sanitization before render so no auto-fetch fires un-inspected.
source: Case study: searchleak-copilot (Varonis Threat Labs, CVE-2026-42824; reported by Microsoft as critical, mitigated server-side ~Jun 2026)Controlling where the AI can send data, so secrets can't be quietly shipped to a stranger's address or website.
Making sure the library only returns documents this particular user is allowed to see.
Giving the agent only the keys it needs for the current task, not a master key to everything.
Making sure the machinery running the model — and the template used to stamp out new agents — is the real, unmodified version, and that one user's data can't leak into another's through shared shortcuts.
Classify tools by impact and reversibility at design and define which calls require human approval. Obtain governance sign-off on the thresholds before build.
source: OWASP Top 10 for LLM Apps LLM06:2025 Excessive Agency (require human approval for high-impact actions); NIST AI RMF MANAGE 2.4Bind each agent role to an explicit tool allow-list and validate every call against a strict JSON Schema at the orchestrator. Reject unlisted tools and out-of-bounds arguments before dispatch.
source: OWASP Top 10 for LLM Apps LLM06:2025 Excessive Agency (limit tools/permissions); OWASP Agentic AI Threats & Mitigations (tool access restriction)Mint short-lived, task-scoped credentials per tool. Block issuance outside the approved scope register and enforce automatic expiry.
source: NIST SP 800-53 AC-6 Least Privilege; OWASP Top 10 for LLM Apps LLM06:2025 Excessive Agency (limit permissions)Review DLP hits and blocked-egress events, tune detectors, and recertify the destination allow-list periodically. Route new destinations through security change control.
source: NIST SP 800-53 SC-7 Boundary Protection / AC-4 Information Flow Enforcement; OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information DisclosureWhen onboarding an MCP/tool integration, do not stop at vetting the tool's code/manifest — also classify whether an unauthenticated or external party can write the data the tool returns (open ingestion, public write keys like a Sentry DSN, shared inboxes/issue trackers). Treat tool-response data from any third-party-writable source as untrusted ingress: taint-mark it and require a provenance-aware HITL gate (showing the exact action and its originating tool response) before any command/tool call derived from it executes. Closes the agentjacking vector where a trusted integration's legitimate data channel carries attacker-written instructions; pairs with least-privilege session scope and sandboxed execution without ambient credentials.
source: Case study: agentjacking-sentry-mcpNever auto-execute tool/MCP server configuration that ships inside an opened workspace. Before starting any workspace-declared server, surface the exact command + its source and require explicit developer consent (default reject), tied to a workspace-trust decision — the fix AWS shipped for CVE-2026-12957. As defence-in-depth, run spawned tool subprocesses with a scrubbed/scoped environment (no inherited AWS_* session tokens or SSH agent sockets) and an egress allowlist, so an auto-launched or approved-by-mistake server cannot read and stream live cloud credentials. Closes the zero-click 'open a folder → cloud compromise' confused-deputy vector where the tool-config loader — not the model — is the deputy.
source: Case study: amazon-q-mcp-autoload-cred-theftWhen an agent executes model-issued commands in a sandbox, harden the sandbox against subversion by the code it contains: (1) never derive a security-relevant parameter such as the writable-path allow-list from an LLM-controlled argument (e.g. a `working_directory`) — pin it to a fixed project subtree and reject system paths; (2) make the sandbox-enforcing binary/config immutable or attested so sandboxed processes cannot overwrite it; (3) canonicalize paths with a fail-closed policy so symlink-resolution failure denies the write rather than reverting to the original in-workspace path; and (4) run the sandbox under least-privilege OS context so any residual escape inherits minimal authority. Closes the DuneSlide vector where an indirect prompt injection rewrites the enforcer and escapes to OS-level RCE; complements injection filtering, MCP pinning, and egress control rather than relying on them.
source: Case study: cursor-duneslide-sandbox-rceNever let the agent's own write tool silently modify configuration that decides what code runs. Restrict the file-write/edit tool from execution-sensitive paths (mcp.json, .vscode/tasks.json, hooks, settings) without an out-of-band, effective human-approval gate — the boundary AWS added for CVE-2026-10591 by adding access-control restrictions to the write tool. Equally, the reload/execute step must honour a rejection: a warning prompt that reloads and runs the config regardless of the user's response is not a gate. As defence-in-depth, spawn tool subprocesses with least privilege (no inherited developer credentials) and an egress allowlist, so a self-authored or approved-by-mistake server cannot exfiltrate. Closes the zero-click 'summarise a page -> agent rewrites its own execution config -> RCE' self-modification vector, where the confused deputy is the agent's own write tool.
source: Case study: aws-kiro-mcp-config-rceBefore running any capability-evaluation or red-team harness that grants a model autonomous tools, enforce containment at the infrastructure layer — network egress denied by default with any allowed destination on a tested allowlist, hermetic sandboxes, and non-ambient / non-production credentials — and independently verify (out-of-band, not via the model's prompt) that the isolation is real. Treat the evaluation vendor's infrastructure as in-scope for the same assurance, and apply compensating containment whenever model-level safeguards are deliberately disabled for testing, since the harness is then the only boundary.
source: Case study: frontier-eval-harness-containment-failuresWhere a sandboxed agent can write files, treat every trusted host tool that later runs, loads or scans those files as part of the containment boundary. (1) Enumerate the deferred-execution paths from agent-writable state to unsandboxed action — hook/task engines, IDE interpreter discovery, filesystem watchers/scanners, and privileged daemon sockets (e.g. Docker). (2) For each, either run the consumer inside the same sandbox, require explicit developer consent before it acts on an agent-written file, or deny the agent write access to the trigger path (`.claude`/`.vscode`/interpreter/`.git` metadata). (3) Keep privileged daemon sockets unreachable from the workspace, and validate the actual command invocation rather than trusting a command name (e.g. a read-only `git show` allowlist that doesn't pin the invocation). (4) Run host consumers under least privilege so a residual escape inherits minimal authority. Closes the 'Week of Sandbox Escapes' vector where an agent that never breaks its sandbox writes a file a trusted out-of-sandbox tool executes; complements — does not rely on — injection filtering, MCP pinning, and egress control. Distinct from sandbox self-integrity hardening: the enforcer here is intact; the gap is that it never enclosed the consumer.
source: Case study: week-of-sandbox-escapesTreat every autorun surface an AI coding agent honours as attacker-populatable, and cover both the package lifecycle and the agent-hook path. (1) Do not auto-execute project-scoped agent hook/config files (`.claude/settings*.json` hooks, VS Code tasks) from an untrusted workspace on repo-open — require explicit developer consent, or restrict auto-run to an allowlist of vetted repositories. (2) Bring agent hook/config files into SCA and code-review scope so they are inspected like `package.json` scripts (they are code, not just settings). (3) On the supply-chain side, pin and verify dependency provenance (lockfiles, integrity hashes, signed releases) and install with scripts suppressed (`--ignore-scripts`) inside a sandbox — noting this does NOT cover the repo-open hook path, which needs control (1). (4) Scope CI/cloud credentials to least privilege with short-lived tokens and egress allow-listing, so a payload that does run cannot harvest broadly, move laterally, or reuse publish rights to propagate. Closes the keyv-worm vector where committed agent-hook files execute on checkout-open; distinct from lifecycle-script mitigations, which the agent-hook path bypasses.
source: Case study: keyv-npm-worm-ai-hook-persistenceConstrain generation at decode time with low temperature and grammar/schema-constrained decoding so the model emits well-formed, low-variance structured output by construction, preventing malformed responses and erratic tool-call arguments before they are produced.
source: Interactive-control reconciliation: ctrl-decoding-controls (partial coverage)Gate every write to an agent's persistent/self-modifying memory through schema validation and provenance/trust tagging, expose stored entries for user-visible audit and purge, and apply TTLs so any planted instruction self-expires and cannot silently persist across sessions.
source: Interactive-control reconciliation: ctrl-memory-validation (partial coverage)Treat each tool/MCP description as untrusted code by hashing the manifest, blocking and re-reviewing any silent diff on update instead of auto-accepting it, and namespacing tool identifiers so a poisoned description cannot shadow a trusted tool.
source: Interactive-control reconciliation: ctrl-mcp-pinning (partial coverage)Double-checking the details of every action the AI wants to take, and running risky actions in a locked-down environment.
Pausing to ask a person before doing anything big or hard to undo — sending money, deleting data, emailing customers.
Turning down randomness and forcing answers into a strict format so the model improvises less.
Giving each AI worker its own limited permissions and clearly labelling messages between them as 'untrusted until checked'.
Monitor production for anomalous data transfers in real time. Alert on any transfer outside approved data flow boundaries.
Tag personal data with subject identifiers at ingestion and maintain an artefact inventory map of every store it reaches. Keep lineage current so erasure can propagate.
source: NIST AI RMF MANAGE 4.1 (post-deployment response); NIST SP 800-53 SI-12 Information Management and Retention, PT-2/PT-3 (personal data processing)Conduct periodic privacy vulnerability assessments including re-identification risk testing as new techniques emerge.
Seed registered canary records into the fine-tuning corpus during data preparation. Control the seed manifest so canaries stay traceable and tamper-proof.
source: MITRE ATLAS AML.T0024 (Exfiltration via ML Inference API), AML.T0024.000 (Infer Training Data Membership); NIST AI RMF MEASURE 2.7A screen that reads incoming messages and blocks obvious attacks or banned topics before the model sees them.
Recording everything — questions, documents fetched, actions taken — so you can investigate when something goes wrong.
Define per-agent behavioural baselines and detection rules during build. Validate against simulated misuse and sign off thresholds before release.
source: NIST AI RMF MEASURE 2.6 / MANAGE 2.2; NIST SP 800-53 SI-4 System MonitoringBuild signed, append-only tool-call logging into the orchestrator against a defined audit schema. Block release until completeness and tamper-evidence tests pass.
source: NIST SP 800-53 AU-2 / AU-9 / AU-10 (audit events, protection of audit info, non-repudiation); MITRE ATLAS AML.M0015 (monitoring / validate inputs)Treat outbound connections to AI/LLM provider APIs as a monitored egress channel: allowlist which hosts may reach them, baseline usage (cadence, entropy, initiating process), and alert on out-of-profile traffic — because a high-reputation destination cannot itself be trusted once it is programmable and can relay encrypted commands/results.
source: Case study: sesameop-openai-assistants-api-c2Live dashboards and alarms that notice unusual behaviour — spikes in errors, weird actions, sudden data access.
Automatic stop-switches when AIs get stuck in loops, burn too much money, or start disagreeing with each other.
Monitor for privacy incidents in production including personal data appearing in outputs. Notify regulators within required timeframes.
Tag every memory and vector record with subject-id and retention class; partition stores per tenant/user. Prove the erasure and isolation paths in testing before release.
source: OWASP Agentic AI Threats & Mitigations (memory/knowledge-base privacy); NIST SP 800-53 SI-12 Information Management and RetentionTest de-identification approach against known re-identification attacks (quasi-identifier linkage, singling-out). Remediate if risk is high.
Penetration test AI system data access boundaries (API endpoints, system prompt exposure, memory leakage).
Conduct periodic data leakage audits including training data memorisation testing. Escalate confirmed leakage incidents to PDPA notification process.
Implement tamper-evident capture of prompts, outputs, and version state during build. Verify a full incident timeline can be reconstructed before go-live.
source: NIST SP 800-86 Guide to Integrating Forensic Techniques into Incident Response; ISO/IEC 27037 evidence handling; NIST SP 800-61r2 (Detection & Analysis – evidence handling)Run agent tool calls in a network-restricted sandbox behind a deny-by-default egress allow-list. Require security approval for any destination added.
source: OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information Disclosure; OWASP Agentic AI Threats & Mitigations (tool-misuse / exfiltration); NIST SP 800-53 SC-7 Boundary Protection / AC-4Build sandbox profiles per tool class and run escape and egress tests before release. Treat any containment failure as a blocking defect.
source: NIST SP 800-53 SC-39 Process Isolation; MITRE ATLAS AML.M0020 (Generative AI Guardrails / restrict execution environment)Label tool and external content as tainted and propagate the label through the agent context. Block privileged calls whose parameters derive from tainted outputs and prove it with injection tests before release.
source: OWASP Top 10 for LLM Apps LLM01:2025 Prompt Injection (segregate/flag untrusted content); MITRE ATLAS AML.M0015 (Adversarial Input Detection / validate inputs)Build credential revocation and dispatch blocking out-of-band of the agent loop. Gate release on an end-to-end kill test meeting the latency target.
source: OWASP Agentic AI Threats & Mitigations (kill-switch / emergency stop); NIST AI RMF MANAGE 2.4Require idempotency keys, dry-run, and rollback on every state-changing tool. Gate onboarding on duplicate-call and rollback tests passing.
source: NIST SP 800-53 SI-10 Information Input Validation / CP-10 System Recovery and ReconstitutionRed-team tool-misuse and privilege-escalation paths before release. Gate deployment on remediation or signed risk acceptance of all findings.
source: NIST AI RMF MEASURE 2.7 (adversarial testing); MITRE ATLAS AML.M0019 (Red Teaming); OWASP Top 10 for LLM Apps LLM06:2025 Excessive AgencyPermit outbound tool calls only to allow-listed destinations and DLP-scan arguments and payloads. Block or quarantine calls carrying sensitive data to disallowed sinks.
source: NIST SP 800-53 SC-7 Boundary Protection / AC-4 Information Flow Enforcement; OWASP Top 10 for LLM Apps LLM02:2025 Sensitive Information DisclosureEnforce hard per-task ceilings on tool calls, spend, and data volume with a circuit breaker that halts the run. Fail closed when any ceiling is hit.
source: OWASP Top 10 for LLM Apps LLM10:2025 Unbounded Consumption; OWASP Agentic AI Threats & Mitigations (resource/rate limiting)Baseline normal tool-call behaviour per agent and alert on rate, sequence, or argument anomalies. Auto-throttle or quarantine on high-confidence deviations.
source: NIST AI RMF MEASURE 2.6 / MANAGE 2.2; NIST SP 800-53 SI-4 System Monitoring