AI & Computational Science

PhantomEnvironments: Training LLM Agents in Fictional Worlds

How the science connects

Reinforcement lear…Language modelSynthetic data

AI Insight

Researchers developed PhantomEnvironments, a system that trains large language model agents using entirely synthetic, rule-generated fictional worlds rather than real-world data. The agents learn to perform multi-hop search tasks by navigating templated articles in fictional universes, acquiring skills that successfully transfer to real-world search benchmarks. The approach eliminates the need for expensive human-curated training data and avoids issues with benchmark contamination, while agents trained on these simple environments often outperform those trained on actual real-world data.


This method provides a cost-effective and scalable alternative for training AI agents without relying on expensive real-world datasets or risking data contamination. The finding that simple rule-based fictional environments can produce agents capable of real-world performance suggests a more efficient pathway for developing capable AI systems across various search and reasoning tasks.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.

Source: PhantomEnvironments: Training LLM Agents in Fictional Worlds