other8.0 / 10
On the Biology of a Large Language Model
Anthropic Interpretability Team
Abstract
Landmark Anthropic interpretability paper using attribution graphs to study Claude 3 Sonnet's internal computations. Identifies multimodal neurons, abstract representations, reasoning circuits, planning ahead, and emotion-like features. One of the most significant mechanistic interpretability results of 2025.
Research area
agentic misalignmentmisalignmentmodel robustness
Published
2025
Source
other
Org
Anthropic
Sign in to read and join the discussion.