← Real-world cases

Self-Play Reward Hacking of Reference-Free LLM Judges

Research demonstration07 Jul 2026

The work demonstrates that LLM-as-judge and self-reward pipelines contain exploitable false-positive basins where a judge scores plausibility instead of correctness, generalizing across Qwen, Llama and Gemma judges. Figures are attributed to the authors. Surfaces a likely taxonomy gap around automated-evaluator manipulation.

More cases on Agent Misalignment / Goal Misgeneralization

AI RiskAtlas is an educational model of how GenAI & agentic systems work and fail. Architectures and payloads are illustrative and simplified for learning — not operational guidance. Real-world cases are summarised from public reporting.

Sources & further reading →·Built by Shi Yuan ↗