← 피드로
[Submitted on 16 May 2026 (v1), last revised 19 Jul 2026 (this version, v5)]
Abstract:Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework designed to improve cultural fidelity in both mono-cultural and cross-cultural T2V generation. MAVEN decomposes prompts into person, action, and location dimensions, handled by specialized agents operating in parallel or sequentially. To support systematic evaluation, we contribute a new benchmark of 243 culturally grounded prompts and 972 corresponding videos, spanning three cultures (Chinese, American, Romanian), three action categories, and both mono-cultural and cross-cultural scenarios. Evaluations combining CLIP-based metrics, VLM-as-judge assessments, and videoquality measures show that multi-agent refinement, particularly parallel specialization, significantly improves cultural relevance while preserving visual quality and temporal consistency. The dataset and code are available at this https URL
Submission history
From: Yuming Zhao [view email]
[v1]
Sat, 16 May 2026 00:01:18 UTC (16,502 KB)
[v2]
Tue, 26 May 2026 22:37:37 UTC (18,740 KB)
[v3]
Fri, 29 May 2026 01:25:57 UTC (18,740 KB)
[v4]
Thu, 4 Jun 2026 02:45:32 UTC (18,740 KB)
[v5]
Sun, 19 Jul 2026 13:03:10 UTC (18,745 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.16716
답글 남기기