Back to papers
arxiv6.0 / 10

Data Attribution at Scale via Influence Matrix Estimation

Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas

Abstract

Data attribution seeks to quantify how individual training examples shape a model's predictions and underpins problems including data valuation, machine unlearning, and model interpretability. Despite having a long line of work, computationally scalable methods often struggle to predict the effect of removing training data in neural networks due to their non-convex nature. To overcome this challenge, metagradient-based methods such as MAGIC (Ilyas and Engstrom, 2025) differentiate each prediction through the entire training run and compute its exact influence with respect to the training data, but require a separate run for every prediction. To reduce this cost, we cast budgeted attribution as estimating a large influence matrix from a small number of measurements. We show that the measurements most appropriate for recovering this matrix differ from those best suited for attribution itself. We then present two algorithms, MAGE and SPELL, suited for reconstruction and attribution respectively, that run on existing metagradient machinery at no extra cost. Empirical studies demonstrate strong performance over existing baselines across training scales and measurement budgets.

Research area

data attributioninterpretabilitymachine unlearning
Published
14 Sept 2026
Source
arxiv
Org
Carnegie Mellon University
View paper
Sign in to read and join the discussion.