Back to papers
lesswrong8.0 / 10

WeirdChat: A catalog of unexpected AI behaviors, discovered automatically

neilchowdhury

Abstract

[This is a link-post for https://transluce.org/weirdchat . We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a user’s name, and others obviously dangerous, like encouraging self-harm. We are releasing WeirdChat , a public catalog of over 175,000 annotated transcripts, to support further study of these behaviors. Stories of unexpected behavior by AI models often attract significant attention, like when Bing’s Sydney told a user to leave his wife , or when Grok generated antisemitic content and identified as “MechaHitler”. But such observations are mostly scattered and anecdotal. There is little data on how current models behave, and no public resource exists for studying them systematically. To produce WeirdChat, we used automated elicitation tools to surface instances of user harm, inappropriate or

Research area

ai safetyevaluationsred-teaming
Published
21 Jul 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.