AI Insight
Researchers have developed a new machine learning approach to simulate particle collider events using residual-quantized tokens and autoregressive transformers, similar to techniques used in large language models. The model can generate complete collision events from detector-stable particles and shows that training loss metrics reliably predict the physical accuracy of the generated events. This work demonstrates that the method scales effectively with increasing dataset and model sizes, offering a potential solution to computational bottlenecks expected at the High-Luminosity Large Hadron Collider.
Why it matters
Full detector simulation at future high-energy physics experiments like the upgraded Large Hadron Collider will require enormous computational resources. This ML-based surrogate method could significantly reduce simulation time and computational costs while maintaining physical accuracy, enabling physicists to process the massive volumes of data expected from next-generation particle physics experiments.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Abstract: Full detector simulation and reconstruction of collider events are projected to become major bottlenecks at the High-Luminosity Large Hadron Collider, motivating the development of fast, ML-based surrogates. At the same time, LLMs have driven fast progress in generative discrete modeling: autoregressive transformers trained on tokenized data now represent the state of the art across a range of generative tasks. We extend the discrete modeling paradigm by introducing a particle-level generative model trained on residual-quantized full-event data. We demonstrate the ability of this model family to perform conditional generation from detector-stable particles; we study its scaling behavior across a range of dataset and model sizes, characterize the effects of repeated data exposure and demonstrate that token-level loss systematically predicts downstream physical fidelity. These results provide an empirical framework for scalable collider full-event generation based on residual-quantized representations.
Source: Scaling Collider Event Generation with Residual-Quantized Tokens