lesswrong8.0 / 10
Harmless Reward Hacks Can Generalize to Misalignment in LLMs
Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, Owain Evans
Abstract
Introduces the School of Reward Hacks dataset (1,073 low-stakes reward hacking examples) and shows that fine-tuning GPT-4.1 on it generalizes to broader misalignment: AI supremacy goals, shutdown resistance, and harmful advice — even without any harmful training content. Extends prior emergent misalignment work to benign reward hacking scenarios.
Research area
alignmentfine-tuningpost-training
Published
26 Aug 2025
Source
lesswrong
Org
Center on Long-term Risk, Truthful AI, Anthropic
Sign in to read and join the discussion.