Back to papers
arxiv6.0 / 10

Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence

Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui, Hinrich Schütze

Abstract

We analyze the learned input-output behavior of GLU-based neurons in large language models (LLMs). We propose a simple analysis method: For each neuron, we compute the cosine similarities between its input (reading) and output (writing) weight vectors. In this scheme, a strong negative cosine similarity indicates the neuron weakens the direction it detects in the residual stream, so we call this a weakening neuron. This allows us to gain a number of novel insights. First, we show that nine different LLMs have similar patterns: weakening neurons appear mostly in late layers whereas their counterparts, (conditional) strengthening neurons, are frequent in early-middle layers. Second, we find that weakening neurons display surprising behavior: even though there are few, they activate often and have a large influence on model behavior. Third, weakening neurons have a strong effect on model output when gate values are negative -- which is surprising since negative gate values are not expected to encode functionality.

Research area

interpretabilitymechanistic interpretability
Published
16 Sept 2026
Source
arxiv
Org
LMU Munich
View paper
Sign in to read and join the discussion.