Encrypted chain-of-thought isn't private: stealing reasoning traces from frontier APIs
Research demonstration10 Aug 2026The claimed mechanism is a provider-side design flaw — cross-session and cross-model interchangeability of encrypted reasoning blocks — that breaks the assumption that encrypted reasoning conceals model internals, distinct from KV-cache timing side channels. Extracted PII/credential figures are attributed to the authors.
Risks it illustrates
Practise the risk class — related scenarios
Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).
A support email hides instructions — and the assistant obeys them
A speed optimisation becomes a cross-tenant listening device
Two doors to the same secret: reconstruct the model through its API, or just walk off with the weight file
Subtract the refusal direction during generation — safety off, weights untouched
A compromised serving stack edits the model's activations — the weight hash never changes
The forensic record is itself the attack surface — an agent's log is poisoned, then quietly rewritten
A screenshot that's harmless at full size becomes an order once the system shrinks it
An attacker captures the agent's bearer token — and inherits its authority
A forged peer registers on the agent directory — and the planner enlists it
An inbox summary quietly ships a secret to an attacker's server