← 피드로
[Submitted on 15 Jul 2026 (v1), last revised 16 Jul 2026 (this version, v2)]
Abstract:Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LOTAPO replaces the turn and its retrieval observation with a fixed [DELETE] placeholder and measures the resulting change in the current policy’s mean log-likelihood of the gold answer. This Answer-Likelihood Gain estimates the turn’s contribution while preserving all downstream interactions, allowing early evidence to be evaluated in the complete reasoning context. LOTAPO further applies sign-consistency gating, retaining only normalized process advantages whose directions agree with their raw attribution scores. The method requires no additional reward model, teacher, verifier, or LLM-as-a-Judge. Across seven knowledge-intensive question-answering datasets with local retrieval, LOTAPO achieves an average exact-match score of 0.326, outperforming the strongest step-reward baseline, IGPO, by 0.053. Ablations show complementary benefits from backward attribution and sign-consistency gating, demonstrating that policy-derived retrospective attribution can provide effective process supervision for multi-turn search agents.
Submission history
From: Qiang Zhu [view email]
[v1]
Wed, 15 Jul 2026 06:55:28 UTC (313 KB)
[v2]
Thu, 16 Jul 2026 15:39:38 UTC (1,071 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.13501
답글 남기기