Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven
Marc Carauleanu
Abstract
This research was conducted at Overlap Research and supported by BlueDot Impact . Summary We tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss . Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct and Gemini 2.5 Pro were deceptive on 96-100% of trials in our main evaluation before fine-tuning. After SOO SFT, deception was reduced to 30.24%, 22.48%, 21.76%, and 6.00%, respectively. A direct instruction for honesty had little effect overall. [1] The less convenient results are also important: Generalization was much stronger across variations of the original scenario than across two more distant scenarios. One model, Gemma-3-27B-It, showed almost no improvement on the extended scenarios. MT-Bench fell by 0.82-1.39 points for the three open-weight models we tested, so this intervention did not preserve capabilities for free. Qualitative Gemini traces suggest that SOO SFT can change which character the model identifies with, but these examples do not establish a mechanism. Our current view is that SOO SFT is a promising, scalable behavioral intervention, no