Diagnosing and Mitigating Context Rot in Long-horizon Search

작성자

카테고리:

← 피드로
arXiv cs.AI · Shijie Xia, Yikun Wang, Zhen Huang, Pengfei Liu · 2026-08-05 AI

[Submitted on 29 Jun 2026 (v1), last revised 4 Aug 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon search tasks. The concern that increasing context length degrades model capabilities, known as context rot, has become a widely recognized issue for these applications. However, in deep search scenarios, it remains unclear how models actually fail under extensive context, and to what extent existing methods can mitigate such failures. Through a systematic study of four flagship models across three benchmarks, we identify a previously overlooked phenomenon, which we term premature termination: under extensive context, models give up or provide uncertain incorrect answers long before exhausting the context window. By controlling for query difficulty, we show that the premature termination rate is positively correlated with context length. Based on the findings, we revisit methods to mitigate context rot, including context management and parallel sampling. For context management, we analyze seven methods across three categories and show that they are inherently test-time scaling strategies that reduce the premature termination rate to enable more exploration, and we further provide model-dependent principles for method selection. For parallel sampling, we develop a behavior-aware filtering strategy and observe a performance gain of 2.6% to 4.9% across three aggregation methods.

Submission history

From: Shijie Xia [view email]
[v1] Mon, 29 Jun 2026 02:53:28 UTC (3,601 KB)
[v2] Tue, 4 Aug 2026 16:37:54 UTC (5,990 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.29718

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다