AI Insight
Researchers tested a deep-learning model (TAF-Net) trained on one Alzheimer's cohort to predict MCI-to-AD conversion in an independent cohort of 101 OASIS-3 participants. The model's risk scores correlated with established biological markers including faster atrophy in AD-signature brain regions, cognitive decline rates, and reduced functional connectivity in key brain networks. However, while the model's ability to rank patients transferred across cohorts, its absolute risk predictions did not calibrate well, and its prognostic signal was largely explained by structural brain atrophy already captured by conventional measures.
Why it matters
This validation demonstrates that AI models for Alzheimer's prediction can capture real biological processes across different patient populations, but highlights critical limitations: the models may not add substantial information beyond traditional brain imaging measures, and their risk probabilities require recalibration for each new clinical setting before use in individual patient counseling.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Background. Deep-learning models predict mild cognitive impairment (MCI)-to-Alzheimer’s disease (AD) conversion from structural MRI with high accuracy. However, they are typically validated within a single cohort and judged on discrimination alone. Whether their risk scores are biologically grounded, and whether they generalise to independent data, remains unclear. Methods. We applied the ADNI-trained Temporal Adaptive Fusion Network (TAF-Net), without retraining, to 101 MCI participants from OASIS-3 (25 converters, 76 stable) and tested whether its conversion-risk scores were grounded in independent structural, functional, and cognitive markers of AD. Regional atrophy rates were derived from longitudinal FreeSurfer, cognition from longitudinal MMSE and CDR-Sum-of-Boxes, and baseline resting-state functional connectivity from fMRIPrep, in structural and functional subsamples of 41 and 40 participants. Associations used rank-based statistics with false-discovery-rate correction. Results. Higher TAF-Net risk tracked faster atrophy in medial-temporal AD-signature regions but not in AD-spared cortex, an anatomically specific coupling that survived adjustment for global atrophy. In external validation, risk discriminated converters (AUC = 0.72), comparable to native atrophy and strongly concordant with it; atrophy statistically accounted for the model’s prognostic signal. Discrimination transferred but calibration did not: the two lowest tertiles of risk were assigned near-zero probability yet converted at 15%. Risk also tracked the multi-year rate of cognitive decline and, at baseline, was associated with reduced within-network functional-connectivity integrity, concentrated in salience and default-mode hubs; longitudinal functional analyses were underpowered. Conclusions. A conversion model trained on one cohort produced risk scores that, in an independent cohort, were grounded in the structural, functional, and cognitive hallmarks of Alzheimer’s disease, supporting biological validity and external generalisation. The score, however, largely re-expresses the neurodegenerative substrate captured by structural atrophy. Rank ordering transferred across cohorts but absolute risk did not, so the score requires recalibration before its values can be interpreted as individual probabilities.