Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

작성자

카테고리:

← 피드로
arXiv cs.AI · Faizan Iqbal · 2026-07-24 AI

[Submitted on 25 May 2026]

View PDF HTML (experimental)

Abstract:We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three categories – multi-sensor joint assessment, response proportionality, and pattern disambiguation – with 1,800 API calls at temperature 0.0, we find that all tested models consistently produced no precautionary warning signal across the tested scenarios where multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor threshold violations. All five models (ChatGPT-4o, Gemini 2.5 Flash, DeepSeek, Kimi, Llama 3.1 8B) score near zero on Category A multi-sensor scenarios (Q2: 0.000-0.208; Q3: 0.000-0.592) compared to strong performance on single-sensor scenarios (Category B Q1: 0.975-1.000). Structured tabular formatting shows no consistent advantage over plain prose; ChatGPT-4o performs significantly better under prose (p = 0.001). These findings have direct implications for practitioners deploying the tested models in physical safety monitoring systems.

Submission history

From: Faizan Iqbal [view email]
[v1] Mon, 25 May 2026 05:43:37 UTC (681 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.20476

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다