HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

작성자

카테고리:

← 피드로
arXiv cs.AI · Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin · 2026-06-29 AI

[Submitted on 19 Mar 2026 (v1), last revised 26 Jun 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows. Within the controlled frame-budget regime that governs practical deployment, prior selectors score frames against a single global query embedding; as a result, compositional multimodal questions that involve temporal ordering or cross-modal cues such as “what happens on screen right after the narrator mentions the reaction?” are flattened into a representation that loses sub-event ordering and modality bindings. We introduce \textbf{HiMu}, a training-free framework for compositional multimodal frame selection. A single text-only LLM call decomposes the query into a hierarchical logic tree whose leaves are atomic predicates, each routed to a lightweight expert spanning vision (CLIP, open-vocabulary detection, OCR) and audio (speech recognition and non-speech sound matching). Expert signals are normalized, smoothed to align across modalities, and composed bottom-up through fuzzy-logic operators that enforce temporal sequencing and adjacency, yielding a continuous per-frame satisfaction curve. Under the standard 16-frame budget on Video-MME, LongVideoBench, and HERBench-Lite, HiMu achieves state-of-the-art accuracy among frame selection methods and improves over uniform sampling across seven diverse MLLMs as a drop-in module, matching the accuracy of uniform sampling at $4\times$ its frame budget, without retraining and without multiple iterative MLLM calls during selection.

Submission history

From: Dan Ben Ami [view email]
[v1] Thu, 19 Mar 2026 07:11:53 UTC (26,841 KB)
[v2] Fri, 26 Jun 2026 08:15:42 UTC (29,262 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2603.18558

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다