← 피드로
[Submitted on 20 Apr 2025 (v1), last revised 4 Aug 2026 (this version, v3)]
Abstract:AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question. When self-play search is given a useful prior, does the network absorb the induced behavior, or does the behavior stay rented from search at test time? We answer with Cross-Phase Prior Intervention (CPI), which switches a root-level tactical prior on and off independently during training and during evaluation, separating the prior’s online effect from the learned residual it leaves in the weights. The endpoint is deliberately narrow: how often a network discharges a forced defensive obligation when no search-time guidance is available. On a sealed one-shot final test in 9×9 Gomoku and 19×19 Go, deleting the prior still leaves a large residual response rises from 13.8% to 26.3% in Gomoku and from 0.6% to 33.8% in Go-and soft reweighting teaches as well as hard action restriction, so pruning legal actions is not the mechanism. The same cross bounds the claim: re-enabling the prior restores nearly 100% response, leaving dependence gaps of 73.7 and 65.8 points. A latched-position evaluation localizes the residual to trained geometry-absent at the shared initialization, emerging over training, worth +17.3 points on in-distribution defenses but only +1.3 on structurally novel ones. Search is therefore best read as a training-time behavioral curriculum whose lessons are real, partial, and geometry-bound, and online competence and internalized competence are different estimands that a diagonal ablation cannot tell apart.
Submission history
From: Binjie Guo [view email]
[v1]
Sun, 20 Apr 2025 14:29:39 UTC (2,031 KB)
[v2]
Mon, 23 Mar 2026 15:02:43 UTC (733 KB)
[v3]
Tue, 4 Aug 2026 10:18:59 UTC (1,167 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2504.14636
답글 남기기