AI Insight
This study introduces "instruction retrieval," a method that enables small language models to solve expert-level problems by retrieving tailored instructions at inference time rather than through fine-tuning. A teacher model creates a reusable corpus of instructions for problem clusters within a domain, containing background knowledge, step-by-step procedures, and common mistakes. Testing on medical, legal, and mathematical benchmarks showed the method improved accuracy substantially, with MedQA accuracy increasing by 10.6 percentage points compared to 3.9 points from retrieved textbook passages.
Why it matters
This approach allows resource-constrained edge devices to perform specialized tasks without expensive fine-tuning for each domain or requiring continuous access to large models. The reusable instruction corpus can democratize access to expert-level AI capabilities across fields like medicine and law while reducing computational costs.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a specific domain or task writes the knowledge into the parameters but must be repeated for every model and domain, and a retrieved passage leaves the model to find the relevant fact and apply it on its own. We introduce instruction retrieval, which distills a teacher model’s expertise into a corpus of instructions tailored so that a small model can follow. For each cluster of a domain’s problems, the teacher writes one instruction with the background knowledge the cluster depends on, a procedure for that kind of problem, and the common mistakes made on it. This reusable corpus needs only to be built once per domain and requires no run-time teacher access. At inference, a frozen small model retrieves the instructions nearest its question and follows them, with no fine-tuning. Across medicine, law, and mathematics benchmarks, we demonstrate the corpus improves over zero-shot on every task and over few-shot prompting, self-consistency, and other retrieved text on medicine and law. On MedQA the corpus raises mean accuracy by 10.6 points, where retrieved textbook passages raise it by 3.9 and few-shot examples from the same teacher lower it. An error analysis shows that only the background knowledge fixes the questions a small model always gets wrong, and that the procedure and common mistakes fix only the questions where it wavers between options. Our results show that automatically retrieved inference-time procedural guidance and domain knowledge can yield substantial gains for small models.
Source: Instruction Retrieval at Inference Time for Small Language Models