Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

작성자

카테고리:

← 피드로
arXiv cs.AI · Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim · 2026-08-31 AI

[Submitted on 29 Sep 2025 (v1), last revised 31 Aug 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes and relational clauses. To address this problem, we propose restructuring linguistic representations according to the hierarchical relations within sentences for language-based object detection. A key insight is that textual tokens should be disentangled into core components-objects, attributes, and relations-and aggregated into hierarchically structured sentence-level representations. Building on this principle, we introduce the TaSe (Talk in Pieces, See in Whole) framework with three main contributions: (1) a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; (2) the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and (3) aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives. Experimental results under the OmniLabel benchmark show a 24% performance improvement, demonstrating the importance of linguistic compositionality.

Submission history

From: Sojung An [view email]
[v1] Mon, 29 Sep 2025 02:14:26 UTC (4,861 KB)
[v2] Fri, 28 Aug 2026 14:49:57 UTC (6,880 KB)
[v3] Mon, 31 Aug 2026 07:36:14 UTC (6,880 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2509.24192