Back to papers
lesswrong8.0 / 10

The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness

Charlie Griffin

Abstract

1) The safe-to-dangerous shift is a fundamental problem for eval realism Suppose we have a capable and potentially scheming model, and before we deploy it, we want some evidence that it won’t do anything catastrophically dangerous once we deploy it. A common approach is to use black-box alignment evaluations. However, alignment evaluations are only reassuring to the extent that the model can't reliably [1] distinguish the deployment distribution from the evaluation distribution, as it is otherwise difficult to rule out the possibility of alignment faking . There are many approaches one could use to try to make evaluations appear more realistic: you can try to create realistic environments (e.g. Petri , WebArena , OSWorld ); use data from past deployments (e.g. OpenAI , SAD ); and spoof tool-call responses (e.g. ToolEmu ). However, the core difference between an alignment evaluation and a real deployment, is that an alignment evaluation must be safe [2] . That is, if you want to safely evaluate an untrusted AI system, you need to limit your AI’s ability to cause harm within your evaluation environment. If you want to deploy your AI system (usefully) you need to give it at least some

Research area

alignmentevaluationsafety and security
Published
14 May 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.