A Research Agenda for Secret Loyalties
Joe Kwon
Abstract
Frontier AI models serve millions of military personnel on classified networks, support operational military targeting, automate scientific pipelines in national laboratories, generate and review significant volumes of production code, and increasingly automate the development of its successors. The more responsibilities AI systems accumulate, the more valuable it becomes for a malicious actor to covertly influence what they do. In a new paper, we define and analyze a specific threat: secret loyalties . We argue that this threat is neglected but tractable, and propose a research agenda to address the technical side. We thank Alan Chan, Aniket Chakravorty, Abbey Chaver, Lukas Finnveden, Tim Hua, Max Nadeau, Aris Richardson, Alexandra Souly, and Anna Wang for helpful feedback and discussions. Full paper: https://www.formationresearch.com/secret-loyalties-whitepaper.pdf X Thread: https://x.com/TomDavidsonX/status/2054614224437907770 What is a secret loyalty? A model has a secret loyalty when it has been intentionally caused to advance a specific actor's interests (which we call the principal ) and this orientation is not disclosed to operators, auditors, users, or other affected parti