Back to papers
lesswrong8.0 / 10

Malign initializations are more robust when the model can think better in the reasoning language than in the output language

Dylan Xu

Abstract

One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations . A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here . Specifically, here’s a fairly typical result from our previous research : We train a (reasoning) malign init to sandbag on some inputs. We SFT the model on responses to simple questions, generated by a different model which is prompted to speak like a pirate. The training in (2) generalizes to updating the malign init’s reasoning to no longer think about sandbagging. Instead, the model solves the problem in its reasoning (while speaking like a pirate) and doesn’t sandbag anymore. Our malign inits’ fragility might be due to feedback spillover : the pirate outputs are not generated by a malign model, and so SFT generalizes to make our malign inits’ reasoning non-malign too. In this post we study a strategy for fixing this issue, and thereby making more robust malign inits; we call our strategy “ dumbspeak ”. The basic idea of the strategy is to assume that the AI model understands some languag

Research area

alignmentdeceptive alignmentevaluations
Published
27 Aug 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.