AI Insight
This study developed and validated genetic prediction models for metabolomic biomarkers in populations of admixed American ancestry, primarily using data from the Mexico City Prospective Study with 132,336 participants. Models trained on this admixed population substantially outperformed those trained on European ancestry populations, achieving median predictive accuracy improvements from R-squared of 0.027 to 0.083. When applied to predict cardiometabolic disease associations in the All of Us cohort, the admixed-ancestry-trained models identified five times more significant associations than European-trained models for conditions including heart disease, type 2 diabetes, and chronic kidney disease.
Why it matters
This research addresses a critical gap in metabolomics research by providing improved tools for studying populations underrepresented in genetic studies, who often bear disproportionate burdens of metabolic diseases. The publicly available genetic scores enable more accurate and equitable biomarker discovery in diverse American populations, potentially improving disease prediction and prevention strategies for communities historically excluded from precision medicine advances.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Despite metabolomics transforming our understanding of risk factors and aetiology of metabolic diseases, profiling is rarely performed for people of non-European ancestries, on whom much of metabolic disease burden falls. Metabolome-wide association studies (MWAS) can be performed using genetic scores to predict metabolomic traits, helping address these inequities; however, their performance in populations of admixed American (AMR) ancestries is unexplored. We evaluated 141 genetic scores, developed in an INTERVAL Study sample of European (EUR) genetic ancestries, in the Mexico City Prospective Study (MCPS; n=132,336), obtaining a median predictive R2 of 0.027. Training Bayesian ridge models within MCPS substantially improved performance, with a median R2 of 0.083 on a withheld 20% subset. MCPS-trained models also outperformed INTERVAL-trained models among UK Biobank participants of AMR ancestries (n=600; median R2: 0.070 vs. 0.046). Finally, among AMR participants of the All of Us cohort, using MCPS-trained (vs. INTERVAL-trained) models to predict metabolomic traits yielded five times as many significant associations (FDR-corrected P<0.05) across three cardiometabolic diseases: ischaemic heart disease, type 2 diabetes, and chronic kidney disease. The genetic scores are openly available at the OmicsPred portal (www.OmicsPred.org), enabling better-powered analyses in diverse AMR cohorts and helping reduce global inequities in omics research.