Biology

AI Benchmark Tests Predictions of Antimicrobial Peptides Against Deadly Bacteria

AI Insight

This study introduces AMPBench-MT, a new benchmark for evaluating computational methods that predict antimicrobial peptide (AMP) properties including potency against specific bacterial species, spectrum of activity, and safety profiles. The research demonstrates that models performing well at simple AMP recognition tasks do not necessarily predict real-world assay results accurately, with protein language model embeddings showing the highest prediction errors for minimum inhibitory concentration values. The benchmark reveals that evaluation metrics must account for endpoint-specific behaviors rather than relying solely on binary classification accuracy.


This work addresses a critical gap in AMP drug discovery by showing that current computational methods may appear successful in lab tests but fail to predict clinically relevant properties like toxicity and species-specific effectiveness. The standardized benchmark enables researchers to better assess which computational approaches will translate to successful antimicrobial development, potentially accelerating the discovery of new antibiotics needed to combat drug-resistant infections.


arXiv:2607.25518v1 Announce Type: cross
Abstract: Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-species potency, hemolysis, toxicity, and selectivity. Existing AMP and peptide benchmarks cover binary recognition, multilabel annotation, assay regression, or broader peptide-model comparison, but they do not jointly place AMP recognition, species-conditioned potency, spectrum, safety-facing proxy endpoints, and cross-endpoint behavior within one sequence-homology-controlled protocol. To address this problem, we introduce AMPBench-MT, a provenance-preserving benchmark that standardizes canonical peptide records and organizes them into binary recognition, species-conditioned pMIC regression, and endpoint-specific potency and safety-facing readouts. Across 161 endpoint-specific model evaluations, high binary performance does not reliably indicate assay-endpoint behavior. Frozen protein-language-model embeddings form the leading pMIC error cluster, while graph and classical regressors remain close. Spectrum labels further reveal that PR-oriented metrics can be misleading under scarce observed negatives, whereas low-toxicity, HC50 hemolysis, and selectivity expose smaller but more assay-facing signals. AMPBench-MT shows that AMP evaluation should move beyond recognition leaderboards toward endpoint-aware evidence auditing. Our proposed benchmark is available at https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.

Source: AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction