AI Insight
Researchers analyzed 298 long-read assembled human genomes and extensive transcriptome data to identify and characterize 2,713 protein-coding genes in duplicated genomic regions that are absent from the standard human reference genome. They found that 60% of these duplicated genes are expressed and maintain functional open reading frames, with nearly half showing high expression in brain, embryo, or testis tissues. The study also reclassified 236 presumed pseudogenes as functional protein-coding genes and identified that about one-quarter of duplicated genes show evidence of evolutionary constraint, suggesting functional importance.
Why it matters
This work reveals a substantial portion of human genetic diversity previously hidden in difficult-to-sequence duplicated regions, which may contribute to individual variation in disease susceptibility, brain development, and reproduction. The findings demonstrate that pangenome approaches can uncover functional genes missed by single-reference genome methods, potentially improving genetic diagnostics and our understanding of human evolution.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.
Source: Pangenome discovery and characterization of human protein-coding duplicated genes