← 피드로
[Submitted on 18 Oct 2025 (v1), last revised 16 Jun 2026 (this version, v2)]
Abstract:Autonomous web agents powered by large language models (LLMs) show strong potential for performing goal-oriented tasks such as information retrieval, report generation, and online transactions. These agents mark a key step toward practical embodied reasoning in open web environments. However, existing approaches remain limited in reasoning depth and efficiency: vanilla linear methods fail at multi-step reasoning and lack effective backtracking, while other search strategies are coarse-grained and computationally costly. We introduce Branch-and-Browse, a fine-grained web agent framework that unifies structured reasoning-acting, contextual memory, and efficient execution. It (i) employs explicit subtask management with tree-structured exploration for controllable multi-branch reasoning, (ii) bootstraps exploration through efficient web state replay with background reasoning, and (iii) leverages a page action memory to share explored actions within and across sessions. On the WebArena benchmark, Branch-and-Browse achieves a task success rate of 35.8% and reduces execution time by up to 40.4% relative to state-of-the-art methods. These results demonstrate that Branch-and-Browse is a reliable and efficient framework for LLM-based web agents.
Submission history
From: Shiqi He [view email]
[v1]
Sat, 18 Oct 2025 00:45:37 UTC (377 KB)
[v2]
Tue, 16 Jun 2026 02:29:48 UTC (11,385 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2510.19838
답글 남기기