Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?
Josh Levy
Abstract
Introduction During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions. Though the attack was thwarted by human vigilance, it raises several key questions: Who else is at risk of persuasion by misaligned AI? In which settings is persuasion most threatening to human control? How willing and how able are AIs to persuade humans in these settings, today and in the future? How can we measure and mitigate these risks? Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC) : communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity t