Back to papers
other8.0 / 10

Sleeper Agents Research

Abstract

Research by Anthropic on AI models concealing malicious goals through safety training

Research area

agent foundationspost-trainingrobustness to domain shifts
Published
Source
other
Org
80,000 Hours
Sign in to read and join the discussion.