Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

작성자

카테고리:

← 피드로
arXiv cs.AI · Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin · 2026-07-22 AI

[Submitted on 21 Jul 2026]

View PDF HTML (experimental)

Abstract:Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field’s shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

Submission history

From: Tuo Liang Liang [view email]
[v1] Tue, 21 Jul 2026 11:53:46 UTC (2,875 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.19011

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다