AI Insight
This study demonstrates that current methods for detecting positive selection in protein evolution produce false positives when they fail to account for natural selection acting on synonymous mutations. The researchers developed BUSTED+S+MSS, an improved statistical framework that incorporates models of synonymous selection, and showed across five diverse taxonomic groups that this correction reduces the number of genes incorrectly identified as evolving under positive selection while improving overall model accuracy. The work reveals that ignoring selection on synonymous sites, driven by factors like translational efficiency and mRNA stability, systematically biases estimates of adaptive evolution.
Why it matters
This research provides a more accurate tool for identifying genes undergoing adaptive evolution, which is crucial for understanding evolutionary processes, predicting disease resistance, and identifying drug targets. The correction is particularly important for highly divergent species comparisons where synonymous selection effects are most pronounced.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
The ratio of nonsynonymous to synonymous substitution rates ({omega}) constitutes a fundamental parameter for inferring adaptive protein evolution, predicated upon the assumption that synonymous substitutions are selectively inert. This premise, however, is increasingly untenable given evidence of selection acting on synonymous substitutions, driven by various biological processes such as translational efficiency and mRNA stability. In this study, we demonstrate that unmodelled synonymous selection introduces substantial bias into {omega} estimation, resulting in elevated false positive rates in tests for positive selection. To rectify this, we present BUSTED+S+MSS, a statistical framework incorporating Multiclass Synonymous Substitution (MSS) models into BUSTED, a method for detecting episodic selection. By partitioning synonymous codons into empirically derived rate classes, this approach accounts for global synonymous constraints. Application to five diverse clades – Drosophila, Caenorhabditis, Enterobacteria, Saccharomyces, and Primates – reveals that the inclusion of MSS components consistently improves model fit and reduces the proportion of genes inferred to be under positive selection. In Enterobacteria, genes retaining significance under the corrected model exhibit weaker constraint on synonymous substitutions (dSs), consistent with the hypothesis that unmodelled purifying selection drives spurious signals of adaptation. Furthermore, an information-theoretic analysis indicates that whilst site-specific variation (SRV) provides the primary correction, global synonymous rate variation (MSS) contributes a distinct second-order correction. In highly divergent alignments, these signals act in concert to improve model fit. The BUSTED+S+MSS framework, especially when coupled with an "error-sink" to absorb alignment artifacts, thus offers a computationally feasible means to disentangle adaptive nonsynonymous substitution from the confounding effects of synonymous constraint.