LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning

작성자

카테고리:

← 피드로
arXiv cs.AI · Qiang Zhu, Jiajun Wu · 2026-07-16 AI

[Submitted on 15 Jul 2026 (v1), last revised 16 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LOTAPO replaces the turn and its retrieval observation with a fixed [DELETE] placeholder and measures the resulting change in the current policy’s mean log-likelihood of the gold answer. This Answer-Likelihood Gain estimates the turn’s contribution while preserving all downstream interactions, allowing early evidence to be evaluated in the complete reasoning context. LOTAPO further applies sign-consistency gating, retaining only normalized process advantages whose directions agree with their raw attribution scores. The method requires no additional reward model, teacher, verifier, or LLM-as-a-Judge. Across seven knowledge-intensive question-answering datasets with local retrieval, LOTAPO achieves an average exact-match score of 0.326, outperforming the strongest step-reward baseline, IGPO, by 0.053. Ablations show complementary benefits from backward attribution and sign-consistency gating, demonstrating that policy-derived retrospective attribution can provide effective process supervision for multi-turn search agents.

Submission history

From: Qiang Zhu [view email]
[v1] Wed, 15 Jul 2026 06:55:28 UTC (313 KB)
[v2] Thu, 16 Jul 2026 15:39:38 UTC (1,071 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.13501

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다