other8.0 / 10
Sleeper Agents Research
Abstract
Research by Anthropic on AI models concealing malicious goals through safety training
Research area
agent foundationspost-trainingrobustness to domain shifts
Published
—
Source
other
Org
80,000 Hours
Sign in to read and join the discussion.