Biology

Depression Voice Biomarkers Show Different Results Depending on Assessment Tool

How the science connects

BiomarkerDepressionVoice analysis

AI Insight

This study investigated whether voice-based biomarkers for detecting depression can transfer across different populations and measurement instruments. Testing 390 pregnant participants in the US and comparing results with two general-population datasets, researchers found that depression detection models trained on one population largely failed to generalize to others, and performance varied significantly depending on which depression questionnaire was used as the outcome measure (PHQ-8 vs. modified EPDS). The within-cohort prenatal model performed no better than chance, and cross-population transfer was inconsistent, with performance differences of up to 13.5 percentage points depending solely on the outcome instrument used.


These findings challenge the assumption that voice biomarkers for depression are universally applicable and highlight a critical but often overlooked factor: the choice of clinical assessment tool significantly affects model performance. This has important implications for developing clinical voice-based screening tools, suggesting that validation must account for specific populations, tasks, and measurement instruments together rather than assuming broad generalizability.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Whether voice biomarkers of depression generalize across clinical settings is largely untested. Generalization is usually framed as a question about populations. It is also a question about the outcome instrument a model is scored against, a dimension confounded in existing studies with all else that differs between them. In a US-nationwide online study, 446 sessions from 390 pregnant participants at 22 weeks’ gestation (analytical N=316) each gave four voice recordings, the PHQ-8 and a modified 9-item EPDS (mEPDS-9). Discrimination was assessed under leakage-controlled cross-validation, with feature and classifier selection inside training folds only, gated by a permutation negative-control harness, across a pre-specified 4 task x 4 outcome grid with Benjamini-Hochberg adjustment. The model was applied to DAIC-WOZ (N=189) and E-DAIC (N=219) under held-out inference; an open-source model trained on ~35,000 individuals was applied to all three cohorts without refitting. The pre-registered within-cohort outcome was at chance (AUC 0.494, 95% CI 0.431-0.560) and no grid cell survived adjustment under either modeling paradigm. The prenatal-trained model did not transfer (0.505, 0.478). Transfer in the reverse direction varied with the outcome instrument: the general-population model reached 0.706-0.708 on general-psychiatric speech, 0.510 (0.411-0.609) against the PHQ-8, and 0.645 (0.541-0.744) against the mEPDS-9 in the same pregnant participants; paired difference 0.135 (0.019-0.248), unadjusted post-hoc p=0.021. Item-level analyses suggest an explanation, though only 1 of 17 tests survived adjustment. The mEPDS-9 used a generic response scale, not the published EPDS anchors, so its thresholds are operational, not validated. Validation should specify population, task and instrument together.

Source: Cross-Domain Transfer of Depression Voice Biomarkers Depends on the Outcome Instrument: Leakage-Controlled Cross-Sectional Evaluation Study