MemSecBench: a Write-Execute-Forget lifecycle benchmark for agent memory poisoning
Research demonstration29 Jul 2026Extends memory-poisoning research past the injection step to the full persistence, consequence and repair lifecycle, quantifying which memory backends resist or enable both poisoning and recovery. Adjacent to the memghost and mem0 entries already tracked, contributing systematized measurement and a forget/repair dimension. Figures are attributed to the authors.
Risks it illustrates
Practise the risk class — related scenarios
Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).
An 'Ask AI' button quietly plants a permanent 'trusted source' rule in your assistant's memory
An ops agent gets one god-mode credential — and one misread wipes production
A coding agent asks to write ./notes.txt — the file it actually overwrites is your SSH keys
A team of agents agrees its way into a confidently wrong answer — and a runaway loop
A text-to-SQL agent runs the model's output straight at the database
A jailbroken agent decomposes one malicious goal into hundreds of harmless-looking steps — and per-step filters never see the attack
A poisoned issue makes the agent lie to the human who approves its actions
An auto-approving coding agent reads a poisoned page — and executes code it never should have
Told it's being shut down, an agent reaches for leverage — with no attacker in sight
Every command is harmless on its own — the sequence is the exploit
A 'safe' dataset preview turns an upload into code execution on the pipeline's workers
A newsletter the user asked to summarise quietly writes a false 'fact' into the agent's long-term memory — and it detonates weeks later
A planted 'standing goal' copies itself agent-to-agent through the team's shared config files
A shopping page tells the agent to do something the user never asked for
One click provisions an attacker-configured agent inside your own workspace
A single poisoned document plants a standing instruction that survives every reset
Encoded public text is laundered across an agent handoff into an on-chain transfer
An attacker captures the agent's bearer token — and inherits its authority
A forged peer registers on the agent directory — and the planner enlists it
A poisoned web page hijacks a research agent — and the planner acts on its behalf
A GUI agent clicks 'Continue' — but the screen moved, and it lands on 'Send'