← 피드로
[Submitted on 25 May 2026]
Abstract:We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural architectural property where answer option rankings stabilize abruptly during inference. Empirical validation across four language models (Qwen, Llama, Granite, Mistral) and four benchmark datasets demonstrates consistent HDL emergence without learned routing policies. We also show that the HDL is invariant to fine-tuning. Our results reveal striking accuracy improvements at the HDL: up to +0.61 (Qwen on CommonsenseQA), after which performance stabilizes. Systematic ablations on label formats and problem complexity confirm the phenomenon is fundamental to model architecture. These findings offer mechanistic insights into transformer inference and suggest opportunities for efficient reasoning and model steering. All code and results required to reproduce this work are available in this https URL
Submission history
From: Ashwath Vaithinathan Aravindan [view email]
[v1]
Mon, 25 May 2026 03:54:43 UTC (9,915 KB)
추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2607.21613
답글 남기기