Biology

Corpusome, a cross-body-site human microbiome corpus for representation learning

How the science connects

Machine learningMicrobiomeMetagenomics

AI Insight

Researchers have created Corpusome, a comprehensive database of 187,546 human microbiome samples from multiple body sites (gut, oral, skin, respiratory, urogenital) designed specifically for machine learning applications. The database integrates data from three major sources and uses a two-tier structure: 22,588 shotgun sequencing samples with detailed species and pathway information, and 164,958 16S sequencing samples providing broader coverage across body sites. The harmonized metadata allows researchers to account for technical variations between studies, with biological signals from different body sites being approximately 2.4 times stronger than technical variation.


This standardized database addresses a major bottleneck in microbiome research by providing the large-scale, multi-site dataset needed to train more generalizable machine learning models. It could accelerate the development of microbiome-based diagnostics and therapeutics by enabling algorithms that work across different body sites and studies rather than being limited to single cohorts.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a harmonized two-tier cross-body-site human microbiome corpus for representation learning: a harmonized corpus of 187,546 human microbiome samples integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. Corpusome follows a two-tier design preserving both functional depth and cross-body-site breadth: a shotgun tier (22,588 samples, 93 studies) with species- and pathway-level profiles, and a 16S tier (164,958 samples, from a full pull of 708 MGnify studies) with genus-level profiles extending coverage to oral, skin, respiratory, and urogenital sites. It spans six body sites and two modalities, with harmonized metadata for batch-aware modelling. Body-site signal exceeds technical/source variance in the 16S tier by approximately 2.4-fold.

Source: Corpusome, a cross-body-site human microbiome corpus for representation learning