AI & Computational Science

AI Language Models Learn Word Associations, Not True World Understanding

How the science connects

Artificial intelli…Natural language p…Word embedding

AI Insight

This study challenges the interpretation that large language models develop internal world models by showing that static word embeddings can achieve comparable performance on decoding tasks. Across four published test cases involving spatial, temporal, pain, and emotion properties, context-insensitive word vectors achieved substantial decoding accuracy (R² = 0.42-0.59 for coordinates and death years, AUC 0.84-0.88 for pain and emotion classification), approaching or matching transformer model performance. Using disambiguated Wikipedia entity embeddings further closed the performance gap, suggesting that much of what appears to be world modeling may actually reflect distributional word associations rather than deeper semantic understanding.


These findings suggest that successful decoding from language model activations does not necessarily demonstrate that models have formed genuine world representations, as similar results can be achieved through simpler statistical associations. This has important implications for how we interpret and evaluate whether AI systems truly understand concepts versus merely capturing co-occurrence patterns in training data.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models go beyond surface text statistics and form internal models of the world. We show that static word embeddings (fixed, context-insensitive representations learned from corpus statistics) of the same or matched stimuli support much of the same decoding. Across four published cases (place, time, pain and emotion), static vectors predict coordinates and year of death (R^2 = 0.42-0.59), separate pain from matched control sentences (held-out AUC 0.85-0.88), and classify twelve emotions in stories written to avoid naming them (AUC 0.84-0.88). Because static embeddings assign each word a single, context-independent vector, these results are a lower bound on what word associations alone can support. The LLMs retain clear advantages on representational tests, and causal and behavioral findings remain outside the scope of the baseline. On the original authors’ entities, where we reproduce their Llama-2 results, the transformer’s advantage lies mostly in placing historical figures in the right century and places in the right country, coarse sorting that richer word associations would be expected to improve; within those groups every representation orders items poorly. Static vectors for disambiguated Wikipedia entities, which carry the associations of a particular place or person rather than of the words in its name, close most of the remaining gap, matching Pythia-2.8B on coordinates and Llama-2-7B on year of death. These results indicate that decodability alone cannot distinguish a representation of a property from information already available in fixed distributional associations.

Source: World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models