디자인에 의한 비디오 이해: 데이터 세트가 비디오 모델을 형성하는 방법

작성자

카테고리:

← 피드로
arXiv cs.AI · Lei Wang, Syuan-Hao Li, Piotr Koniusz, Yongsheng Gao · 2026-06-09 AI

[Submitted on 11 Sep 2025 (v1), last revised 8 Jun 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model families, they provide limited insight into why particular architectures emerged and succeeded. In this survey, we argue that the evolution of video understanding is fundamentally shaped by dataset structure. We present a dataset-centric perspective that connects dataset structure, inductive biases, and architectural design within a unified framework. We show that different datasets require models to capture specific invariances and capabilities, such as robustness to viewpoint changes, sensitivity to temporal ordering, reasoning over long-range dependencies, relational interactions, and cross-modal alignment. These requirements naturally give rise to inductive biases, i.e., architectural assumptions that favor particular patterns of reasoning and generalization. From this perspective, milestone architectures, including two-stream networks, 3D CNNs, temporal models, transformers, graph-based methods, and multimodal foundation models, can be understood as architectural responses to the challenges posed by evolving datasets. Building on this framework, we systematically analyze how dataset characteristics have shaped architectural innovation across video understanding tasks and discuss the representational biases induced by different data regimes. By unifying datasets, inductive biases, and architectures into a coherent perspective, this survey offers both a retrospective explanation of the field’s evolution and a forward-looking roadmap toward general-purpose video understanding systems. Code and dynamic video visualizations of dataset-induced biases are available at this https URL.

Submission history

From: Lei Wang [view email]
[v1] Thu, 11 Sep 2025 05:06:30 UTC (10,094 KB)
[v2] Mon, 8 Jun 2026 07:16:30 UTC (29,238 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2509.09151

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다