AI Insight
This study examines whether deep learning models for peptide sequencing rely on actual spectral data or learned sequence patterns from training databases. Researchers developed the Prior Bias Index (PBI) framework and DeNovo-PBI benchmark to test models using modified spectra and synthetic peptide sequences, finding that models exhibit strong preferences for database-derived amino acid patterns even when distinguishing spectral evidence is removed. The work reveals that current models may depend heavily on learned sequence priors rather than purely physical spectral information.
Why it matters
This research highlights potential reliability issues in AI-driven peptide identification tools used in proteomics research and drug discovery. Understanding and quantifying these biases is essential for improving model robustness, especially when analyzing novel or non-canonical peptides that differ from typical database sequences.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Deep learning models have advanced de novo peptide sequencing, but their predictions may reflect both physics-based spectral evidence and learned peptide-sequence priors. Systematically measuring such prior-associated behavior is important for benchmarking model robustness beyond conventional proteomics data. Here, we introduce the Prior Bias Index (PBI), a general framework for measuring the extent to which model behavior shifts toward prior-associated reference patterns under controlled conditions, and implement it as DeNovo-PBI, a benchmark for quantifying prior bias in de novo peptide sequencing models. DeNovo-PBI combines benchmark dataset construction, in silico sequence and spectral perturbation workflows, PBI-based metrics, and analysis algorithms to evaluate three forms of prior-associated behavior: sequence-distribution dependence, database amino-acid-pair order preference, and mutation-group prediction consistency under shared sequence context. In addition to experimentally acquired peptide spectra, we generated in silico spectra from random, natural, and mutated peptide sequences and selectively removed fragment ions that distinguish N-terminal residue orders. Across these assays, deep learning models showed peptide-sequence-distribution-dependent performance and strong directional amino-acid-pair order preferences even when order-diagnostic spectral evidence was removed. DeNovo-PBI provides a quantitative benchmark for measuring, comparing, and interpreting learned bias in de novo peptide sequencing models.