ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning

작성자

카테고리:

← 피드로
arXiv cs.AI · Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng · 2026-08-03 AI

[Submitted on 26 May 2026]

View PDF HTML (experimental)

Abstract:Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compression or test-time control, but the absence of a reusable intermediate interface that can replace discarded history and support continued solving. We further identify a key failure mode of outcome-reward-driven long-chain reinforcement learning: when the model has not solved the task before the window is nearly exhausted, the final-answer reward encourages premature guessing rather than continued careful reasoning. We propose ThinkReset, a text-space instantiation of this view. ThinkReset explicitly constructs reusable intermediate interfaces through interface writeback and reset, and directly optimizes post-reset continuation success. Across multiple long-horizon reasoning benchmarks, this perspective consistently improves success rates under fixed context windows.

Submission history

From: Zijian Zeng [view email]
[v1] Tue, 26 May 2026 02:25:09 UTC (47 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.28642

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다