DocMaster: A Hierarchical Structure-Aware System for Document Analysis

작성자

카테고리:

← 피드로
arXiv cs.AI · Ziqi Chen, Yingli Zhou, Fangyuan Zhang, Quanqing Xu, Chuanhui Yang, Yixiang Fang · 2026-07-11 AI

[Submitted on 9 Jul 2026]

View PDF HTML (experimental)

Abstract:Leveraging large language models (LLMs) to analyze complex documents — such as academic papers, technical manuals, and financial reports — has emerged as a mainstream and critical task in both research and industry. In practice, users must first filter relevant documents from large collections and then conduct in-depth analysis (e.g. question answering) over the selected subset, yet existing systems flatten documents into plain-text chunks, discarding the rich hierarchical structures (sections, tables, figures, equations) and degrading downstream performance. We present DocMaster, a hierarchical structure-aware document analysis system. DocMaster parses documents into hierarchical document trees preserving original layouts and constructs a structure-aware semantic index that enables accurate document filtering and in-depth analysis. We demonstrate DocMaster through an interactive web interface that enables users to upload document collections, construct tree-based and multi-view semantic indices, filter relevant documents via natural-language conditions, and perform follow-up question answering over the filtered results. The source code, data, and demo are available at this https URL.

Submission history

From: Ziqi Chen [view email]
[v1] Thu, 9 Jul 2026 14:33:47 UTC (739 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.08539

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다