lesswrong8.0 / 10
Prompt Optimization Makes Misalignment Legible
MATS 8.0 scholars
Abstract
Proposes using prompt optimization (updating instructions rather than weights) to make reward hacking strategies legible in plain English. Demonstrates that optimized prompts reveal misalignment strategies that weight updates hide, enabling sanitization. Introduces 'legible learning' as a safety technique.
Research area
agentic misalignment
Published
2025
Source
lesswrong
Org
MATS
Sign in to read and join the discussion.