Physics

Chess engines miss outcome differences that favor one player over another

AI Insight

This study analyzed over 16 million chess games from positions that the Stockfish 18 engine evaluated as essentially equal (within 10 centipawns of zero). Despite engine assessments showing no advantage for either side, human players consistently achieved skewed outcomes in specific positions, with some positions reliably favoring White or Black across different player groups, time periods, and rating ranges. The reproducibility coefficient was 0.69 overall and reached 0.94 for the most popular positions, indicating that human-level difficulty differs systematically from engine evaluation even in objectively equal positions.


This finding challenges the assumption that chess engine evaluations fully capture position quality for human players. It has implications for opening preparation, chess education, and AI-assisted decision making in domains where human and machine performance diverge, suggesting that optimal strategies may differ between artificial and human intelligence even when formal evaluations appear identical.


Understand the Science

Artificial intelligence 233 articles Explore Concept → Game theory Concept coming soon Chess engines Concept coming soon

arXiv:2607.25655v1 Announce Type: cross
Abstract: Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games’ actual results and what the players’ ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions — disjoint player-account sets (primary), time, and disjoint rating bands — and on an out-of-sample month eight months later. On the primary split, each position’s skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope’s value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|delta| approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.

Source: Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions