Back to papers
lesswrong8.0 / 10

Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction

Nelson Guda

Abstract

The directions nearest a model's prediction decide what kind of answer you get. The next ones out decide where it goes — about five tokens later. Preprint: https://arxiv.org/abs/2608.12447 ; Supplementary materials ; Code . What do we mean when we say that a transformer model has privileged geometry? I honestly wasn't sure about that when I started down this rabbit hole, because that wasn't the initial point of the work. If you want to jump right to the most unexpected finding, scroll down to #8 where I show temporal stratification of output behavior in residual stream geometry. I began the work that led to this paper with the intention of trying to better understand the undifferentiated bulk of the residual stream that is not directly connected to the prediction mechanism. This was my original motivation for looking at the geometry after removing the prediction direction. Around the same time, I learned about thin-shell geometry, and realized there was likely a connection with residual stream geometry, and the rabbit hole suddenly got deeper. Much of the work in the paper was born from a process of getting experimental findings that don't align with expectations and then figuring

Research area

activation patchinglatent space geometrymechanistic interpretability
Published
21 Aug 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.