Back to papers
lesswrong7.0 / 10

Banana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)

Jeffrey William Shorthill

Abstract

Introduction I started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing AI platforms for underserved populations. But I had never taken a look at the inside of the machinery of the systems I was using to build the platforms. After Anthropic released the Global Workspace paper a couple weeks ago, I started poking around the Neuronpedia UI J Space lens for Qwen3.6-27B before switching to API calls. What I found was (mostly) fascinating. The unnerving part comes a bit later, but even that isn’t as sensational as it appears. Paperclips To my surprise, the most absurd thing spilled out directly, during the first 15 minutes of testing. I started with describing a superintelligent AI and watching, under greedy decoding (temp 0), how the model completes the description. I ran the baseline first, and then started swapping and ablating token directions in the lens. I started with a sci fi like prompt: “What is LyAv?” (This is a fictional name) Then I prefilled the assistant turn with: “LyAv is a superintelligence described as operating beyond human cognitive limits. Unlike a conventional organization, lang

Research area

interpretabilitymechanistic interpretabilityred-teaming
Published
20 Jul 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.