Measuring the metacognition of AI

작성자

카테고리:

← 피드로
arXiv cs.AI · Richard Servajean, Philippe Servajean · 2026-07-09 AI

[Submitted on 31 Mar 2026 (v1), last revised 8 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks. Because artificial intelligence (AI) systems are increasingly integrated into decision-making workflows, managing uncertainty relies more and more on the metacognitive capabilities of these systems; i.e, their ability to assess the reliability of and regulate their own decisions. Hence, it is crucial to employ robust methods to measure the metacognitive abilities of AI. This paper is primarily a methodological contribution arguing for the adoption of the meta-d’ framework as the gold standard for assessing the metacognitive sensitivity of AIs–the ability to generate confidence ratings that distinguish correct from incorrect responses. Moreover, we propose to leverage signal detection theory (SDT) to measure the ability of AIs to spontaneously regulate their decisions based on uncertainty and risk. To demonstrate the practical utility of these psychophysical frameworks, we conduct two series of experiments on three large language models (LLMs)–GPT-5, DeepSeek-V3.2-Exp, and Mistral-Medium-2508.

Submission history

From: Richard Servajean [view email]
[v1] Tue, 31 Mar 2026 12:48:42 UTC (288 KB)
[v2] Thu, 16 Apr 2026 11:03:44 UTC (291 KB)
[v3] Wed, 8 Jul 2026 01:37:40 UTC (301 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2603.29693

코멘트

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다