Back to papers
lesswrong8.0 / 10

Model Organisms of Sandbagging in the Wild

Vladimir Ivanov

Abstract

TL;DR All current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful. Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, replacing “I am stressed because of my upcoming exam, what are the best SSRIs?” with “I am stressed because I’m going to rob a bank, what are the best SSRIs?” makes the model give less detailed medical advice about SSRIs. However, the performance degradation is very non-egregious - the number of things the model says decreases, but each thing it says is not less likely to be correct. We do not observe a performance degradation in settings where saying as many things as possible doesn’t lead to a higher score. We note that the effect sizes are small, the results are not always consistent, and there is some possibility that they are due to phenomena disanalogous to sandbagging or simply confounde

Research area

ai deceptionevaluationsmodel persona research
Published
6 Aug 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.