← 피드로
[Submitted on 13 Feb 2026 (v1), last revised 12 Aug 2026 (this version, v2)]
Abstract:Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently evaluated. This work examines the performance of state-of-the-art LLMs and machine translation systems in crisis-domain translation, with a focus on preserving urgency, a critical property for effective crisis communication and triage. Using multilingual crisis data (TICO-19, 30 languages) and a newly introduced urgency-annotated dataset of 100 scenarios translated into 29 languages, we show that dedicated translation models and LLMs exhibit substantial quality degradation, particularly for low-resource languages. Beyond translation quality, we conduct a human annotation study revealing a striking asymmetry: human assessors maintain consistent urgency judgments regardless of prompt language, while LLM-based urgency classifications vary widely across languages for identical scenarios, at times spanning the full range from Not Urgent to Critical. These findings highlight significant risks in deploying general-purpose language technologies for crisis triage and underscore the need for multilingual, human-centered evaluation frameworks.
Submission history
From: Belu Ticona [view email]
[v1]
Fri, 13 Feb 2026 20:56:06 UTC (640 KB)
[v2]
Wed, 12 Aug 2026 06:37:32 UTC (2,670 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2602.13452
답글 남기기
댓글을 달기 위해서는 로그인해야합니다.