AI & Computational Science

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer’s Screening

How the science connects

Alzheimer's diseaseMedical imagingKnowledge transfer

AI Insight

Researchers developed MINT, a framework that transfers knowledge from MRI brain scans to speech analysis for detecting mild cognitive impairment (MCI) in early Alzheimer's disease. The system trains a speech classifier using structural information learned from MRI scans, but only requires speech input during actual screening, eliminating the need for expensive imaging equipment. Testing on the ADNI-4 dataset showed that the speech-based approach performed comparably to standard speech classifiers, while combining both modalities improved upon MRI-only classification.


This approach could enable large-scale, cost-effective screening for early Alzheimer's disease using only speech analysis, while maintaining biological grounding from neuroimaging research. By eliminating the need for MRI scans during screening, this method could make early detection accessible to broader populations, particularly in settings where imaging infrastructure is limited or unavailable.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Alzheimer’s disease is a progressive neurodegenerative disorder in which mild cognitive impairment (MCI) precedes dementia. Structural MRI provides biomarkers but requires costly infrastructure, limiting population-scale deployment. Speech offers a non-invasive alternative, yet speech-only classifiers are developed independently of neuroimaging and lack biological grounding for CN-versus-MCI classification. We propose MINT (Multimodal Imaging-to-Speech Knowledge Transfer), a three-stage framework that transfers MRI-derived biomarker structure to speech during training. An MRI teacher defines a compact embedding space for CN-versus-MCI classification, while a residual projection head aligns speech representations to this space using a combined geometric loss. The frozen MRI classifier enables imaging-free inference. On ADNI-4, aligned speech achieves performance comparable to speech baselines, while multimodal fusion improves over MRI alone. Ablations identify dropout regularization and self-supervised pretraining as important design choices. To our knowledge, MINT is the first demonstration of MRI-to-speech knowledge transfer for early Alzheimer’s screening without imaging at inference.

Source: MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening