AI & Computational Science

Scaling Automatic Research Agents via World Models

How the science connects

Reinforcement lear…Autonomous agentWorld model

AI Insight

Researchers developed World Model RL (WMRL), a new training approach for autonomous research agents that addresses a critical scalability bottleneck in reinforcement learning. The method replaces costly real-time environment execution with a world model simulation, incorporating statistical techniques to handle model imperfections through bias correction and noise reduction. The approach achieves 3-4x training speedup while enabling smaller 4B and 9B parameter agents to outperform much larger 48B and 120B parameter models on research automation benchmarks.


This work directly advances the automation of scientific research by making AI research agents more computationally efficient to train. The method's transferability to other domains like embodied robotics policies suggests broader applications for scaling AI systems that learn through environmental interaction.


Understand the Science

Reinforcement learning 43 articles Explore Concept → Autonomous agent Concept coming soon World model Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we propose World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck. Additionally, the world model can be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations, Online Debiasing and Inverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. Beyond AutoResearch, WMRL also transfers to post-training embodied VLA policies, which demonstrates the generalizability of our method.

Source: Scaling Automatic Research Agents via World Models