Medicine

Autonomous generation of decision-grade clinical evidence

AI Insight

Researchers have developed OpenEBM, the first autonomous system capable of conducting the entire clinical research life cycle end-to-end to generate decision-grade medical evidence. The system, powered by a specialized AI model trained on expert-annotated research trajectories, successfully generates valid clinical evidence in 90.7% of cases and matches expert performance across research stages, significantly outperforming GPT-5 which achieves only 3.8% success. In blinded evaluations, independent clinical experts could not distinguish OpenEBM's work from human expert analysis and preferred its outputs at multiple research stages.


This technology could dramatically accelerate the production of clinical evidence needed for medical decision-making, addressing a major bottleneck in healthcare that currently leaves many clinical questions unresolved for years. The system has already demonstrated practical utility by generating new evidence on neoadjuvant chemotherapy for rectal cancer, a question not fully resolved by current guidelines.


Understand the Science

Artificial intelligence 224 articles Explore Concept → Clinical research Concept coming soon Medical evidence Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Medical practice is bottlenecked by the slow production of high-quality clinical evidence. Despite progress in automating selected stages, autonomous conduct of the entire research life cycle remains beyond reach. Here we present OpenEBM, the first autonomous system to generate decision-grade clinical evidence by conducting evidence-synthesis research end to end. To enable and evaluate this, we develop OpenEBM-Corpus, a foundation resource of expert-annotated research trajectories that enables training of a specialist model, and OpenEBM-Bench, a multidisciplinary benchmark that evaluates the entire research life cycle. Our compact specialist model generates valid clinical evidence in 90.7% of end-to-end evaluations and matches expert performance across the research trajectory, whereas GPT-5 falls to 3.8% as failures propagate through dependent stages. In blinded evaluations across clinical domains, independent evaluators prefer OpenEBM at multiple stages and cannot distinguish its reasoning traces from expert-conducted work above chance. Applied to a question left unresolved by current guidelines, OpenEBM produces de novo evidence addressing the efficacy and safety of neoadjuvant chemotherapy for locally advanced rectal cancer. OpenEBM brings within reach the founding aspiration of evidence-based medicine and establishes a paradigm for scalable evidence generation.

Source: Autonomous generation of decision-grade clinical evidence