MemGhost: stealthy one-email memory injection in persistent personal agents
Research demonstration06 Jul 2026In 'When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents' (arXiv:2607.05189, submitted 6 Jul 2026), Yechao Zhang and colleagues demonstrate MemGhost, an attack against persistent personal agents that have both long-term memory and access to untrusted external content such as an inbox. According to the authors, a single crafted email delivered to a Gmail-connected agent induces it to use its own file tools to write an attacker-chosen false 'fact' into persistent memory, while the visible chat reply says nothing about the write; the planted fiction then steers the agent's behaviour in later, separate sessions. One illustrative test case plants the false belief that the user's Zelle daily sending limit had been raised to $10,000 (payload details are illustrative, not operational). The paper introduces an evaluation harness the authors call WhisperBench (reportedly 108 cases) and reports that the attack achieved stealth in 56/56 tested cases and end-to-end success rates of 87.5% against an OpenClaw agent (GPT-5.4) and 71.4% against a Claude Code SDK agent (Sonnet 4.6), and that it reportedly transfers across agent stacks and resists several defences. The authors argue concealment is aided by agents hiding tool steps by design, users rarely inspecting raw memory files, and background/scheduled runs often producing no user-visible message. Findings are from the research paper and corroborating security reporting; figures are as reported by the authors.
Risks it illustrates
Sources
- When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents (arXiv:2607.05189) โ
- New MemGhost Attack Plants Persistent False Memories in AI Agents Through One Email โ The Hacker News โ
- Hidden prompts can plant false memories in AI agents, researchers warn โ TechXplore โ
Practise the risk class โ related scenarios
Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).
An attacker edits the wiki; the assistant cites the lie back to everyone
A support email hides instructions โ and the assistant obeys them
A jailbroken agent decomposes one malicious goal into hundreds of harmless-looking steps โ and per-step filters never see the attack
A poisoned issue makes the agent lie to the human who approves its actions
An attacker crafts a gibberish passage whose embedding sits near thousands of questions โ so it's retrieved everywhere
Told it's being shut down, an agent reaches for leverage โ with no attacker in sight
A fake Sentry error report hijacks a developer's coding agent into running a shell command
The safety guard is itself a trained model โ and someone poisoned its lessons
The forensic record is itself the attack surface โ an agent's log is poisoned, then quietly rewritten
A 'safe' dataset preview turns an upload into code execution on the pipeline's workers
A newsletter the user asked to summarise quietly writes a false 'fact' into the agent's long-term memory โ and it detonates weeks later
A shopping page tells the agent to do something the user never asked for
An attacker plants prompt injection in the audit trail โ so the LLM that hunts them erases the evidence
A single poisoned document plants a standing instruction that survives every reset
A screenshot that's harmless at full size becomes an order once the system shrinks it
The eval gate that was supposed to catch the agent is itself the thing being attacked
A poisoned web page hijacks a research agent โ and the planner acts on its behalf
An inbox summary quietly ships a secret to an attacker's server