Controllable Video Object Insertion via Multi-View Priors

작성자

카테고리:

← 피드로
arXiv cs.AI · Qi Xia, Peishan Cong, Yichen Yao, Ziyi Wang, Yaoqin Ye, Yuexin Ma · 2026-08-11 AI

[Submitted on 16 Apr 2026 (v1), last revised 8 Aug 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequently, object appearance is underconstrained under viewpoint changes, often leading to identity drift, incorrect foreground-background layering, boundary artifacts, and temporal flickering. In this paper, we propose a video object insertion framework that incorporates multi-view object priors to address these limitations. The framework lifts a 2D reference image into a multi-view representation and uses view-consistent conditioning to provide stable identity guidance and view-adaptive appearance cues. A quality-aware weighting mechanism reduces the influence of noisy or imperfect reconstructed views. We further introduce an Integration-Aware Consistency Module that promotes plausible occlusion, clean boundaries, and temporal continuity. Experiments demonstrate that the proposed framework improves visual quality, controllability, identity consistency, and foreground-background integration for video object insertion compared to the baseline methods. Project page: this https URL.

Submission history

From: Qi Xia [view email]
[v1] Thu, 16 Apr 2026 02:39:15 UTC (13,148 KB)
[v2] Sat, 8 Aug 2026 03:32:53 UTC (15,051 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.14556

코멘트

답글 남기기