AI Insight
CompoWorld introduces a method for training AI agents by automatically generating complex tasks that require using multiple tools and services together. The system creates 448 services with over 10,000 tools, then generates training tasks where information must flow between different services, combining supervised learning from verified task solutions with reinforcement learning guided by completion-focused rewards. When applied to the Qwen3.6-35B-A3B model, this approach improved performance by 9.17 points on average across eight benchmarks and exceeded frontier models like Claude Opus 4.6 on AutomationBench.
Why it matters
This work addresses a key limitation in AI agent training by moving beyond single-environment tasks to multi-service workflows that better reflect real-world automation needs. The compositional approach could enable more capable AI assistants that can chain together multiple tools and services to accomplish complex, multi-step objectives.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Abstract: Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Source: CompoWorld: Compositional Environment Scaling for General Agents