Medicine

AI Dentistry Tool Shows Promise But Makes Predictable Diagnostic Errors

How the science connects

Artificial intelli…Medical diagnosisDentistry

AI Insight

This study evaluated six multimodal large language models on 50 image-based periodontal diagnostic questions, finding that diagnostic errors arose from two mechanistically distinct sources: perceptual failures in extracting visual features from images and cognitive failures in clinical reasoning despite accurate perception. When models were provided with expert-corrected visual descriptions for questions they initially answered incorrectly, some improved their diagnoses while others persisted in error, demonstrating that aggregate accuracy scores mask fundamentally different failure modes with different implications for clinical deployment. The researchers established a dual-process error taxonomy adapted from diagnostic reasoning theory to systematically classify whether errors originated at the visual perception stage or the clinical reasoning stage.


This framework provides a generalizable method for evaluating medical AI systems beyond simple accuracy metrics, enabling developers and clinicians to identify whether failures stem from image interpretation or clinical knowledge gaps. Understanding these distinct error sources is crucial for responsible deployment of AI diagnostic tools in dental education and clinical practice, and for targeting specific model improvements in perceptual versus reasoning capabilities.


Understand the Science

Artificial intelligence 344 articles Explore Concept → Medical diagnosis Concept coming soon Dentistry Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Multimodal large language models (MLLMs) are increasingly applied to image-based clinical reasoning, yet their diagnostic reliability in periodontal image interpretation, and the underlying source of their errors, remain poorly characterized. This study evaluated six architecturally distinct MLLMs (Claude Sonnet 4.5, GPT-5.0, Gemini 2.5, GLM-4.6, Sonar, and Grok 4.1) using 50 image-based multiple-choice questions drawn from the American Academy of Periodontology In-Service Examination, spanning clinical photographs, histopathology, radiographs, cardiac rhythm strips, and anatomical illustrations. A sequential two-phase experimental design was used: in Phase 1, each model independently described each image, selected an answer, and provided a supporting citation; in Phase 2, applied only to questions answered incorrectly, models were given an expert-validated visual description and asked to re-answer, allowing diagnostic improvement through visual correction to be measured directly. Expert ground truth for image content was established by a board-certified periodontist and independently validated by a second board-certified periodontist. Model outputs were classified using a dual-process error taxonomy adapted from Norman’s model of diagnostic reasoning, distinguishing perceptual errors, arising from inaccurate visual feature extraction, from cognitive errors, arising from flawed reasoning despite accurate perception, with cognitive errors further subdivided into correctable and persistent subtypes, and additional categories capturing compound perceptual-cognitive failures and compensatory reasoning that overcame inaccurate perception. Diagnostic accuracy and error type distribution varied significantly across models and image modality. Correcting inaccurate visual descriptions in Phase 2 improved diagnostic accuracy for a subset of previously incorrect responses, indicating that a meaningful share of errors originated at the level of visual perception rather than clinical reasoning; conversely, a distinct subset of errors persisted despite accurate corrected visual input, indicating reasoning-level failures independent of perceptual accuracy. Some models also reached correct answers despite generating inaccurate image descriptions, reflecting compensatory reasoning resilient to perceptual error. These findings show that aggregate accuracy scores conflate mechanistically distinct failure modes, and that perceptual and cognitive errors carry different implications for how MLLMs might be safely deployed or improved for diagnostic image interpretation. The expert-guided visual correction framework introduced here provides a generalizable, mechanism-based approach to benchmarking multimodal AI diagnostic performance that extends beyond periodontics to other visually driven diagnostic domains in medicine. As MLLMs become increasingly accessible to clinicians, residents, and dental educators, distinguishing perceptual from cognitive failure is essential for guiding responsible clinical use, targeting model refinement, and informing AI-augmented dental education and competency assessment.

Source: Expert-Guided Visual Correction for Characterizing Diagnostic Performance and Error Patterns of Multimodal Large Language Models Using Periodontal In-Service Examination Images