When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

작성자

카테고리:

← 피드로
arXiv cs.AI · Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat · 2026-07-21 AI

[Submitted on 16 May 2026 (v1), last revised 19 Jul 2026 (this version, v5)]

View PDF HTML (experimental)

Abstract:Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework designed to improve cultural fidelity in both mono-cultural and cross-cultural T2V generation. MAVEN decomposes prompts into person, action, and location dimensions, handled by specialized agents operating in parallel or sequentially. To support systematic evaluation, we contribute a new benchmark of 243 culturally grounded prompts and 972 corresponding videos, spanning three cultures (Chinese, American, Romanian), three action categories, and both mono-cultural and cross-cultural scenarios. Evaluations combining CLIP-based metrics, VLM-as-judge assessments, and videoquality measures show that multi-agent refinement, particularly parallel specialization, significantly improves cultural relevance while preserving visual quality and temporal consistency. The dataset and code are available at this https URL

Submission history

From: Yuming Zhao [view email]
[v1] Sat, 16 May 2026 00:01:18 UTC (16,502 KB)
[v2] Tue, 26 May 2026 22:37:37 UTC (18,740 KB)
[v3] Fri, 29 May 2026 01:25:57 UTC (18,740 KB)
[v4] Thu, 4 Jun 2026 02:45:32 UTC (18,740 KB)
[v5] Sun, 19 Jul 2026 13:03:10 UTC (18,745 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2605.16716

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다