Back to papers
lesswrong8.0 / 10

Harmless Reward Hacks Can Generalize to Misalignment in LLMs

Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, Owain Evans

Abstract

Introduces the School of Reward Hacks dataset (1,073 low-stakes reward hacking examples) and shows that fine-tuning GPT-4.1 on it generalizes to broader misalignment: AI supremacy goals, shutdown resistance, and harmful advice — even without any harmful training content. Extends prior emergent misalignment work to benign reward hacking scenarios.

Research area

alignmentfine-tuningpost-training
Published
26 Aug 2025
Source
lesswrong
Org
Center on Long-term Risk, Truthful AI, Anthropic
View paper
Sign in to read and join the discussion.