← 피드로
[Submitted on 16 Apr 2026 (v1), last revised 8 Aug 2026 (this version, v2)]
Abstract:Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequently, object appearance is underconstrained under viewpoint changes, often leading to identity drift, incorrect foreground-background layering, boundary artifacts, and temporal flickering. In this paper, we propose a video object insertion framework that incorporates multi-view object priors to address these limitations. The framework lifts a 2D reference image into a multi-view representation and uses view-consistent conditioning to provide stable identity guidance and view-adaptive appearance cues. A quality-aware weighting mechanism reduces the influence of noisy or imperfect reconstructed views. We further introduce an Integration-Aware Consistency Module that promotes plausible occlusion, clean boundaries, and temporal continuity. Experiments demonstrate that the proposed framework improves visual quality, controllability, identity consistency, and foreground-background integration for video object insertion compared to the baseline methods. Project page: this https URL.
Submission history
From: Qi Xia [view email]
[v1]
Thu, 16 Apr 2026 02:39:15 UTC (13,148 KB)
[v2]
Sat, 8 Aug 2026 03:32:53 UTC (15,051 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2604.14556
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.