Back to papers
arxiv8.0 / 10

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi

Abstract

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.

Research area

interpretabilityjailbreakingmechanistic interpretability
Published
4 Sept 2026
Source
arxiv
Org
Thales
View paper
Sign in to read and join the discussion.