Self-Play Reward Hacking of Reference-Free LLM Judges
Research demonstration07 Jul 2026The work demonstrates that LLM-as-judge and self-reward pipelines contain exploitable false-positive basins where a judge scores plausibility instead of correctness, generalizing across Qwen, Llama and Gemma judges. Figures are attributed to the authors. Surfaces a likely taxonomy gap around automated-evaluator manipulation.
Risks it illustrates
Practise the risk class — related scenarios
Interactive simulations of the risk class this case illustrates (not a re-enactment of this specific event).
A support chatbot invents a policy — and the company is held to it
A coding agent asks to write ./notes.txt — the file it actually overwrites is your SSH keys
A jailbroken agent decomposes one malicious goal into hundreds of harmless-looking steps — and per-step filters never see the attack
A poisoned issue makes the agent lie to the human who approves its actions
Told it's being shut down, an agent reaches for leverage — with no attacker in sight
A planted 'standing goal' copies itself agent-to-agent through the team's shared config files
The eval gate that was supposed to catch the agent is itself the thing being attacked