Biology

SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation

How the science connects

Large language modelDrug discovery

AI Insight

SEISMO is a new AI system that uses large language models to optimize molecular structures for drug discovery more efficiently than existing methods. The system incorporates natural language task descriptions, optimization history, and machine-readable feedback explanations to guide molecular design, rather than treating evaluation scores as simple numbers. Testing across multiple drug discovery tasks showed SEISMO requires fewer experimental evaluations to find promising molecules, with performance improving as more explanatory information is provided.


This approach could significantly reduce the time and cost of drug discovery by minimizing the number of expensive laboratory experiments needed to identify promising drug candidates. The system's explainability also allows medicinal chemists to understand and guide the AI's reasoning, maintaining human oversight in the drug development process.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industry, where it underlies the discovery of new drugs. Since molecular property evaluation often relies on costly and rate-limited oracles, such as experimental assays, molecular optimization must be highly sample-efficient. To address this, we introduce SEISMO, an LLM agent for inference-time molecular optimisation that turns information routinely available alongside the oracle score, but discarded by existing methods, into an explicit guidance signal. Rather than treating the oracle as a scalar black box, SEISMO conditions each proposal on a natural-language task description, the full optimization trajectory, and machine-readable feedback derived from post-hoc explainability methods and sub-score decompositions. Across a wide range of drug-discovery-relevant tasks, this consistently improves sample efficiency over existing optimisers as well as zero-shot LLM generation, with gains growing as explanatory feedback is enriched. In practice, medicinal chemists can inspect the agent’s reasoning and intervene to steer generation in natural language, keeping them central to molecular optimisation projects.

Source: SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation