← 피드로
[Submitted on 13 Sep 2026 (v1), last revised 26 Sep 2026 (this version, v3)]
Abstract:Large language model agents are increasingly deployed for long-horizon task execution, raising a central granularity question for trajectory evaluation: whole-trajectory verification is too coarse to capture concrete failures and their associated evidence in long trajectories, while atomic-step scoring is too fine-grained, noise-sensitive, and computationally expensive. This granularity gap makes a single-reference trajectory paradigm inadequate for assessing the rich space of valid agent execution paths and delays timely feedback and early stopping in long-horizon tasks. To address these issues, we propose DynSTEER, a dynamic stage-wise framework for agent trajectory evaluation. DynSTEER bridges the granularity gap through stage-wise dynamic evaluation that segments rollouts at key execution nodes and adapts its multi-level review strategy based on stage-level results; it compiles a path-tolerant milestone graph from available task inputs to preserve diverse legal paths without reference leakage; and it supports terminating unrecoverable agent executions to curb resource waste. Experimental results show that DynSTEER improves evaluation discriminability by over 85\% compared with whole-trajectory evaluation and saves 17.74\% of execution steps. The code is available at this https URL
Submission history
From: Zhichao Shi [view email]
[v1]
Sun, 13 Sep 2026 16:18:03 UTC (652 KB)
[v2]
Tue, 15 Sep 2026 10:27:42 UTC (376 KB)
[v3]
Sat, 26 Sep 2026 08:43:44 UTC (1,562 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.14637