Physics

AI Discovers Scientific Laws Faster by Exploring Before Committing

How the science connects

Machine learningScientific methodPhysics

AI Insight

This study presents a protocol for discovering scientific laws using language models through an "explore-then-commit" approach, where AI proposes hypotheses, plans measurements, and synthesizes governing equations from observations. Testing across 576 trials on physics problems, the method demonstrated significant measurement efficiency improvements, reducing required measurements by approximately 60-80% while substantially decreasing prediction errors. However, improvements in symbolic accuracy were inconsistent across different physics modules, and the study could not definitively attribute measurement savings to specific protocol components.


This approach could accelerate scientific discovery by reducing the number of expensive or time-consuming experiments needed to identify physical laws. The measurement-efficiency gains are particularly relevant for fields where data collection is costly, though the method's effectiveness may vary depending on the type of scientific problem.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Scientific law discovery requires selecting measurements and converting evidence into a governing equation. We evaluate an explore-then-commit protocol in which a large language model proposes hypotheses, a programmatic planner gathers measurements, and a fresh prompt synthesizes the final law from fixed observations. The protocol combines structured probes, automatic numerical diagnostics, restricted measurement batches, and optional interpreter access. Across 576 NewtonBench trials, we compare eight configurations on 12 physics modules using GPT-4.1-mini and a medium-difficulty GPT-4.1 replication. On medium tasks, interpreter-enabled planners use 8.6 versus 22.5 measurements per trial for GPT-4.1-mini and 8.9 versus 43.0 for GPT-4.1. Their mean magnitude-based root-mean-squared logarithmic error falls from 2.514 to 0.202 and from 0.626 to 0.149, respectively. An additional audit retains incomplete and invalid submissions in a coverage-sensitive analysis. Observed symbolic-accuracy gains are less consistent across modules, and random acquisition is competitive with disagreement scoring. Measurement savings occur in every module, but unequal batch constraints prevent attributing them solely to acquisition quality. These results support the complete protocol as a promising measurement-efficient configuration, while leaving its causal components and generalization beyond noiseless direct-equation tasks unresolved.

Source: Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models