Interdisciplinary

Bacterial Species Classification Fails When Based on Genetic Sequences Alone

How the science connects

DNA barcodingBacterial taxonomyMolecular phylogen…

AI Insight

Researchers analyzed 6,660 Rhizobium bacterial 16S rRNA gene sequences from GenBank and found that identical genetic sequences frequently corresponded to different species names, host plants, and geographic locations. Among sequence pairs sharing 100% identity, 66.59% differed in all three metadata categories, while only 1.40% matched across species, host, and country. The study also introduced a novel "microbial h-index" concept to quantify sequence recurrence in databases, finding an h-index of 201 for Rhizobium, meaning the 201st-ranked sequence had 202 identical matches.


This challenges the widespread practice of using 16S rRNA gene sequences alone for bacterial identification and classification, particularly in environmental DNA studies. The findings emphasize the need for comprehensive taxonomic approaches combining genetic, phenotypic, and ecological data to accurately capture bacterial diversity and function.


Understand the Science

DNA barcoding Concept coming soon Bacterial taxonomy Concept coming soon Molecular phylogenetics Concept coming soon

by Rosella Muresu, Monica Rodriguez, Andrea Squartini

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were ‘citations.’ The more one finds the same query sequence, the more that sequence can be considered biologically ‘cited’, i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon’s sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

Source: GenBank mining reveals novel insights into <i>Rhizobium phylogeny</i>: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct ‘microbial h-index’