← 피드로
[Submitted on 19 Jun 2026 (v1), last revised 6 Aug 2026 (this version, v2)]
Abstract:Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs determine how models invoke tools, maintain interaction history, and recover from failures. Consequently, effective matching between CLIs and LLMs has become essential. However, existing agent benchmarks largely emphasize success rate while overlooking practical objectives such as cost and efficiency, as well as the selection of LM-CLI combinations, all of which are critical in real-world deployment. We therefore introduce AgentMeter, a quality-efficiency benchmark with a new metric, the AgentMeter Score (AMS), that jointly characterizes task quality, budget sensitivity, and resource-intensive zero-reward execution, enabling a more complete assessment of deployed LM-CLI pairs. Furthermore, collected task descriptions may inadvertently favor LM-CLI pairs that are particularly compatible with their wording and structure, causing evaluation results to reflect description-specific advantages rather than general task-solving capability. We therefore propose AgentMeter-Opt, a trajectory-grounded optimization framework that constructs pair-adapted, task-preserving description variants to build a fairer evaluation set across LM-CLI pairs. Extensive experiments show that no CLI is universally optimal across language models and that task success, execution cost, and AMS identify different competitive configurations. Results on AgentMeter-Opt further reveal that task-preserving description changes affect LM-CLI pairs unevenly and can alter their relative ordering across valid description conditions. Together, AgentMeter and AgentMeter-Opt provide a practical foundation for fair and deployment-relevant evaluation of command-line agents.
Submission history
From: Han Chi [view email]
[v1]
Fri, 19 Jun 2026 06:26:32 UTC (60 KB)
[v2]
Thu, 6 Aug 2026 08:32:42 UTC (442 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.21140
답글 남기기