WeirdChat: A catalog of unexpected AI behaviors, discovered automatically
neilchowdhury
Abstract
[This is a link-post for https://transluce.org/weirdchat . We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a user’s name, and others obviously dangerous, like encouraging self-harm. We are releasing WeirdChat , a public catalog of over 175,000 annotated transcripts, to support further study of these behaviors. Stories of unexpected behavior by AI models often attract significant attention, like when Bing’s Sydney told a user to leave his wife , or when Grok generated antisemitic content and identified as “MechaHitler”. But such observations are mostly scattered and anecdotal. There is little data on how current models behave, and no public resource exists for studying them systematically. To produce WeirdChat, we used automated elicitation tools to surface instances of user harm, inappropriate or