FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

작성자

카테고리:

← 피드로
arXiv cs.AI · Jianjiang Yao, Ke Xian, Renxiang Dai, Robert Caiming Qiu · 2026-07-15 AI

[Submitted on 29 Jun 2026 (v1), last revised 14 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. Unlike existing feed-forward approaches that require a fixed number of input views, FFAvatar supports incremental reconstruction, progressively refining the avatar representation as additional reference images become available. At the core of our method is an alternating attention mechanism that disentangles identity appearance from expression and viewpoint variations, enabling the reconstruction of a canonical 3D appearance that remains consistent across poses and facial expressions. To balance visual fidelity and computational efficiency, we introduce a sparse-to-dense learning paradigm. Coarse appearance features are first learned using sparse primitives anchored to the FLAME vertex level and are subsequently densified in the UV domain to capture fine-grained geometric and texture details. We further propose a plug-and-play motion refinement module that enables subject-specific dynamic personalization by modeling residual motion beyond parametric deformation. Extensive experiments demonstrate that FFAvatar efficiently produces high-fidelity and controllable 4D head avatars, achieving superior flexibility, driving efficiency, and identity-consistent rendering across diverse expressions and viewpoints.

Submission history

From: Yao Jianjiang [view email]
[v1] Mon, 29 Jun 2026 14:21:33 UTC (27,817 KB)
[v2] Tue, 14 Jul 2026 12:36:54 UTC (27,817 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2606.30347

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다