In-Run Data Shapley for Adam Optimizer

작성자

카테고리:

← 피드로
arXiv cs.AI · Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu · 2026-07-23 AI

[Submitted on 30 Jan 2026 (v1), last revised 22 Jul 2026 (this version, v4)]

View PDF HTML (experimental)

Abstract:Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard. While recent “In-Run” methods bypass the prohibitive cost of retraining by estimating contributions dynamically, they heavily rely on the linear structure of Stochastic Gradient Descent (SGD) and fail to capture the complex dynamics of adaptive optimizers like Adam. In this work, we demonstrate that data attribution is inherently optimizer-dependent: we show that SGD-based proxies diverge significantly from true contributions under Adam (Pearson $R approx 0.11$), rendering them ineffective for modern training pipelines. To bridge this gap, we propose Adam-Aware In-Run Data Shapley. We derive a closed-form approximation that restores additivity by redefining utility under a fixed-state assumption and enable scalable computation via a novel Linearized Ghost Approximation. This technique linearizes the variance-dependent scaling term, allowing us to compute pairwise gradient dot-products without materializing per-sample gradients. Extensive experiments show that our method achieves near-perfect fidelity to ground-truth marginal contributions ($R > 0.99$) while retaining $sim$95% of standard training throughput. Furthermore, our Adam-aware attribution significantly outperforms SGD-based baselines in data attribution downstream tasks.

Submission history

From: Meng Ding [view email]
[v1] Fri, 30 Jan 2026 21:31:40 UTC (2,415 KB)
[v2] Fri, 6 Feb 2026 15:27:57 UTC (2,416 KB)
[v3] Sun, 8 Mar 2026 18:56:18 UTC (2,421 KB)
[v4] Wed, 22 Jul 2026 02:36:11 UTC (1,622 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2602.00329

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다