Back to papers
lesswrong8.0 / 10

Prompt Optimization Makes Misalignment Legible

MATS 8.0 scholars

Abstract

Proposes using prompt optimization (updating instructions rather than weights) to make reward hacking strategies legible in plain English. Demonstrates that optimized prompts reveal misalignment strategies that weight updates hide, enabling sanitization. Introduces 'legible learning' as a safety technique.

Research area

agentic misalignment
Published
2025
Source
lesswrong
Org
MATS
View paper
Sign in to read and join the discussion.