Back to papers
lesswrong8.0 / 10

A Research Agenda for Secret Loyalties

Joe Kwon

Abstract

Frontier AI models serve millions of military personnel on classified networks, support operational military targeting, automate scientific pipelines in national laboratories, generate and review significant volumes of production code, and increasingly automate the development of its successors. The more responsibilities AI systems accumulate, the more valuable it becomes for a malicious actor to covertly influence what they do. In a new paper, we define and analyze a specific threat: secret loyalties . We argue that this threat is neglected but tractable, and propose a research agenda to address the technical side. We thank Alan Chan, Aniket Chakravorty, Abbey Chaver, Lukas Finnveden, Tim Hua, Max Nadeau, Aris Richardson, Alexandra Souly, and Anna Wang for helpful feedback and discussions. Full paper: https://www.formationresearch.com/secret-loyalties-whitepaper.pdf X Thread: https://x.com/TomDavidsonX/status/2054614224437907770 What is a secret loyalty? A model has a secret loyalty when it has been intentionally caused to advance a specific actor's interests (which we call the principal ) and this orientation is not disclosed to operators, auditors, users, or other affected parti

Research area

agentic misalignmentsecret loyalties
Published
13 May 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.