Back to papers
lesswrong7.0 / 10

Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas

oakhu

Abstract

Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. [1] We give an initial empirical demonstration of this effect on Kimi K2.6. The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. [2] To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations. We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize. Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input. Background Suppose that you're a language model in a prisoner's dilem

Research area

decision theoryemergent risk & game-theoretic analysismulti-agent systems
Published
15 Aug 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.