Selective Off-Policy Reference Tuning with Plan Guidance

작성자

카테고리:

← 피드로
arXiv cs.AI · Anh Duc, Tien-Phat Nguyen, Thien Huu Nguyen, Linh Ngo Van, Trung Le · 2026-09-28 AI

[Submitted on 12 May 2026 (v1), last revised 24 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models.

Submission history

From: Tien-Phat Nguyen [view email]
[v1] Tue, 12 May 2026 04:25:41 UTC (7,881 KB)
[v2] Wed, 13 May 2026 06:41:19 UTC (7,864 KB)
[v3] Thu, 24 Sep 2026 19:01:19 UTC (7,865 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.11505