Back to papers
other8.0 / 10

On the Biology of a Large Language Model

Anthropic Interpretability Team

Abstract

Landmark Anthropic interpretability paper using attribution graphs to study Claude 3 Sonnet's internal computations. Identifies multimodal neurons, abstract representations, reasoning circuits, planning ahead, and emotion-like features. One of the most significant mechanistic interpretability results of 2025.

Research area

agentic misalignmentmisalignmentmodel robustness
Published
2025
Source
other
Org
Anthropic
View paper
Sign in to read and join the discussion.