Back to papers
other8.0 / 10

Stress Testing Deliberative Alignment for Anti-Scheming Training

Abstract

Partnership with OpenAI to assess frontier language models for early signs of scheming in controlled stress-tests and study training methods to reduce these behaviors

Research area

agentic misalignment
Published
Source
other
Org
Apollo Research
Sign in to read and join the discussion.