AI & Computational Science

AI Drug Discovery Models Fail When Tested on Novel Molecular Structures

How the science connects

Machine learningDrug discovery

AI Insight

This study demonstrates that AI models for predicting molecular properties fail significantly when tested on structurally novel compounds that are physicochemically distant from training data. Using a new "structural-frontier" evaluation method on six drug absorption, distribution, metabolism, excretion, and toxicity (ADMET) tasks, researchers found that prediction errors increased by a median of 87% and mean of 130% compared to standard scaffold-based splitting methods. Advanced graph neural networks and specialized training penalties did not resolve these failures, suggesting fundamental limitations in current approaches to predicting properties of unfamiliar molecular structures.


These findings reveal critical weaknesses in AI drug discovery tools when encountering novel chemical structures, which is precisely the scenario where such tools are most needed for discovering new therapeutics. The research highlights that current evaluation methods may give falsely optimistic estimates of model performance in real-world drug development applications.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most physicochemically remote scaffold groups, and evaluate it on six public experimental or curated ADMET tasks. Against a 70/10/20 scaffold control with identical acyclic grouping, the frontier inflates equally weighted primary error with a taskwise median of 87.0% and a skew-sensitive mean of 130.3% (descriptive task/seed bootstrap interval, 52.1-246.0%). The mean falls to 75.9% once BBB is removed; that endpoint is the one whose score ranking inverts at the frontier. A message-passing graph-network control still shows a large gap (mean 82.8% over four tasks) and does not invert, so a low-capacity head does not explain the effect. We also test Multi-View Frontier Risk Extrapolation (MV-FREX), a count-adjusted tail-risk penalty over four molecular views, and treat it as a falsifiable probe. It changes normalized frontier error by only 0.16% relative to empirical risk minimization for the perceptron head (interval, -0.43-0.84%) and by -1.9% for the graph network; three fixed robust-penalty controls are likewise inconclusive. Against the published Lo-Hi and DataSAIL splitters, the frontier inflates error more on average, though no split is uniformly hardest. An audit of 31,561 marine natural products further shows that OOD status and agreement with legacy ADMET predictions depend on the molecular view, endpoint, and teacher coverage. Split construction and label provenance are important evaluation constraints in their own right, and the tested training penalties do not resolve the frontier failures we observe.

Source: Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models