AI & Computational Science

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

How the science connects

Artificial intelli…Computer visionMedical imaging

AI Insight

Researchers developed FetalMind, an AI system specifically designed to interpret fetal ultrasound images and generate medical reports. The system uses a novel approach called Salient Epistemic Disentanglement that incorporates expert medical knowledge to better understand relationships between different ultrasound views and potential diseases. Trained on FetalSigma-1M, a new dataset of 20,000 fetal ultrasound reports from twelve medical centers, FetalMind demonstrated 14% average performance improvement over existing methods and 61.2% higher accuracy in detecting critical conditions across all pregnancy stages.


This advancement could improve prenatal care by providing more accurate and consistent interpretation of fetal ultrasound images, particularly important given the complexity and variability of fetal imaging compared to standard adult medical scans. The system's ability to handle multiple viewing angles and rare conditions may assist healthcare providers in resource-limited settings or enhance diagnostic confidence for complex cases.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which poses challenges of multi-view image reasoning, numerous diseases, and image diversity. To bridge this gap, we introduce FetalMind, a medical AI system tailored to fetal ultrasound for both report generation and diagnosis. Guided by clinical workflow, we propose Salient Epistemic Disentanglement (SED), which injects an expert-curated bipartite graph into the model to decouple view-disease associations and to steer preference selection along clinically faithful steps via reinforcement learning. This design mitigates variability across diseases and heterogeneity across views, reducing learning bottlenecks while aligning the model’s inference with obstetric practice. To train FetalMind at scale, we curate FetalSigma-1M dataset, the first large-scale fetal ultrasound report corpus, comprising 20K reports from twelve medical centers, addressing the scarcity of domain data. Extensive experiments show that FetalMind outperforms open- and closed-source baselines across all gestational stages, achieving +14% average gains and +61.2% higher accuracy on critical conditions while remaining efficient, stable, and scalable. Project Page: https://hexiao0275.github.io/FetalMind.

Source: Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation