공간-주파수-광학 흐름 다중 모드 기능의 융합을 기반으로 하는 AIGC 비디오 감지

작성자

카테고리:

← 피드로
arXiv cs.AI · S. Hong, X. Q. Wang, C. Zhang, J. C. Wang, P. X. Duan, Y. W. Wang · 2026-09-23 AI

[Submitted on 17 Aug 2026]

View PDF HTML (experimental)

Abstract:The rapid evolution of generative AI (e.g., Sora, Hunyuan) makes it essential to develop effective detection strategies that can generalize across ever-evolving synthesis techniques. This study is motivated by the observation of a fundamental challenge in generative models: the inherent difficulty of maintaining cross-modal consistency between appearance and motion. To this end, we propose a multi-modal framework for AIGC video forgery detection tasks, named Cross-Attention based Video Forgery Detector (CrossAtt-VFD), based on joint multi-view analysis of this http URL, we introduce a dual-branch architecture that simultaneously extracts spatial-frequency and optical-flow this http URL approach enables the modeling of videos from complementary perceptual this http URL core of this process is a dedicated cross-attention mechanism, which governs the alignment of the two modalities and translates cross-modal inconsistencies into a potent diagnostic signal. This multi-modal strategy facilitates the detection of motion that is statistically inconsistent with the visual appearance of a scene. Comprehensive experimental results demonstrated that our model achieves an accuracy of 94.22%, a precision of 91.67 %,and a recall of 96.25 %, effectively verifying the advantages of the multi-modal fusion strategy.

Submission history

From: Sheng Hong [view email]
[v1] Mon, 17 Aug 2026 05:38:08 UTC (15,457 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2609.26274