Instruction Retrieval at Inference Time for Small Language Models

작성자

카테고리:

← 피드로
arXiv cs.AI · Kenan Alkiek, David Jurgens, Vinod Vydiswaran · 2026-10-01 AI

[Submitted on 15 Oct 2025 (v1), last revised 29 Sep 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a specific domain or task writes the knowledge into the parameters but must be repeated for every model and domain, and a retrieved passage leaves the model to find the relevant fact and apply it on its own. We introduce instruction retrieval, which distills a teacher model’s expertise into a corpus of instructions tailored so that a small model can follow. For each cluster of a domain’s problems, the teacher writes one instruction with the background knowledge the cluster depends on, a procedure for that kind of problem, and the common mistakes made on it. This reusable corpus needs only to be built once per domain and requires no run-time teacher access. At inference, a frozen small model retrieves the instructions nearest its question and follows them, with no fine-tuning. Across medicine, law, and mathematics benchmarks, we demonstrate the corpus improves over zero-shot on every task and over few-shot prompting, self-consistency, and other retrieved text on medicine and law. On MedQA the corpus raises mean accuracy by 10.6 points, where retrieved textbook passages raise it by 3.9 and few-shot examples from the same teacher lower it. An error analysis shows that only the background knowledge fixes the questions a small model always gets wrong, and that the procedure and common mistakes fix only the questions where it wavers between options. Our results show that automatically retrieved inference-time procedural guidance and domain knowledge can yield substantial gains for small models.

Submission history

From: Kenan Alkiek [view email]
[v1] Wed, 15 Oct 2025 15:51:13 UTC (231 KB)
[v2] Wed, 7 Jan 2026 14:13:11 UTC (238 KB)
[v3] Tue, 29 Sep 2026 16:45:12 UTC (120 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2510.13935