← 피드로
[Submitted on 1 Jul 2026 (v1), last revised 4 Jul 2026 (this version, v2)]
Abstract:Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Existing code summarization solutions often rely on a single language model or coding assistant like Claude Code, and treat source code as flat text, underutilizing the rich interdependencies and hierarchical information within a repository. To address these shortcomings, we propose Agent4cs – a multi-agent framework that summarizes large codebases in a bottom-up fashion, where a summarization agent focuses on producing robust summaries; a keyword-extraction agent proactively identifies critical information from subfolders; and a quality-assurance agent iteratively refines the outputs for readability, coherence, and completeness. Evaluated on 7 frontier models, Agent4cs improves semantic consistency across all folder levels by average 8% compared to two structured prompting baselines with code segments. Furthermore, extensive evaluation on real-world datasets demonstrates up to 38% gains in normalized keyword coverage rate over the same baselines.
Submission history
From: Yongjian Tang [view email]
[v1]
Wed, 1 Jul 2026 19:41:38 UTC (1,880 KB)
[v2]
Sat, 4 Jul 2026 12:11:09 UTC (1,880 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.01425
답글 남기기