Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers

작성자

카테고리:

← 피드로
arXiv cs.AI · Zixuan Gong, Shijia Li, Yong Liu, Jiaye Teng · 2026-07-14 AI

[Submitted on 28 Feb 2025 (v1), last revised 11 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However, existing theoretical analyses hardly account for this feature-level two-stage phenomenon, which could be conceptually attributed to disentangled two-type features like syntax and semantics. In this paper, we theoretically demonstrate how the two-stage training dynamics potentially occur in transformers. Specifically, we analyze the feature learning dynamics induced by the aforementioned disentangled two-type feature structure, grounding our analysis in a simplified yet illustrative setting that comprises normalized ReLU self-attention and structured data. Such disentanglement of feature structure is general in practice, e.g., natural languages contain syntax and semantics, and proteins contain primary and secondary structures. To our best knowledge, this is the first rigorous result regarding a feature-level two-stage optimization process in transformers within this theoretical framework. A corollary further indicates that such a two-stage process is closely related to the spectral properties of attention weights.

Submission history

From: Zixuan Gong [view email]
[v1] Fri, 28 Feb 2025 03:27:24 UTC (2,470 KB)
[v2] Sat, 11 Oct 2025 04:45:15 UTC (3,382 KB)
[v3] Sat, 11 Jul 2026 09:47:59 UTC (2,507 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2502.20681

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다