Medicine

AI language models tested on blood disorder medicine questions

How the science connects

Artificial intelli…Medical diagnosisHematology

AI Insight

Researchers evaluated ten large language models on 1,477 board-style hematology multiple-choice questions covering nine disease areas and six clinical skill domains. The top-performing models achieved over 90% accuracy on text-based questions and 75-79% on multimodal questions, with Claude Opus 5 scoring highest at 92.7% and 76.9% respectively. Model accuracy correlated with size, open-weight models showed larger improvements between generations than proprietary models, and top performers displayed similar error patterns on difficult cases.


This study provides the first comprehensive benchmark of AI performance in specialist hematology, addressing critical questions about whether LLMs can safely assist with specialist-level medical queries. The findings suggest these models possess substantial hematology knowledge but still require expert oversight, informing how they might be integrated into clinical practice and medical education.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.

Source: Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology