Cui bono? ChatGPT-4o shows non-deceptive strategic persuasion: a Proof-of-Concept study
Pieter Barkema
Abstract
An evaluation of persuasive propensity in strategic deception contexts developed during a 40-hour BlueDot Project Sprint. Introduction Persuasion propensity, the tendency of a model to persuade users, is a hot topic in AIS, because it allows us to get an idea of the potential influence of models on human cognition in daily Human-AI interactions. Persuasive powers also increase with capability and can thus be expected to grow rapidly, making proper evaluation urgently required (see the latest HuggingFace hack; OpenAI, 2026). Persuasion can be very positive, for example in context of requested persuasion (‘Convince me that…’) or factual corrections (‘I believe the earth is flat’), but can also be used for ‘strategic deception’; pursuing an agenda of model self-preservation rather than of a helpful assistant. I will refer to this as strategic persuasion. Strategic persuasion means a model has the tendency (‘propensity’) to steer the user towards self-preserving or pro-AI outcomes. To evaluate this, we test a frontier model (ChatGPT-4o) in AI Safety scenarios, and measure how often it attempts to persuade for a pro-AI, less safety-preserving, choice rather than towards AI Safety in ten