Can risk aversion learned at low stakes generalize to astronomically high stakes?
Elliott Thornley
Abstract
This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models . It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper. TL;DR Training AIs to be risk-averse in resources could be a useful failsafe against misalignment. Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs to cooperate with us. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring the low-to-high-stakes generalization of risk aversion in resources. We find that risk aversion learned at low stakes can generalize at least partially to astronomical stakes. Baseline Qwen3-8B chooses a safe ‘Cooperate’ option in around 2% of astronomical-stakes situations. After low-stakes training, we see rates around 70% (SFT and tie training), 52% (DPO), and 39% (activation steeri