Biology

AI models mirror how human brains align thinking and action during gameplay

How the science connects

Reinforcement lear…Functional magneti…

AI Insight

This study examined how vision-language models (VLMs) and large-action models (LAMs) align with human brain activity during video game playing, using fMRI data from participants playing Atari-style games. Both model types showed better alignment with brain activity than traditional reinforcement learning agents, with improvements being most pronounced in higher-order frontal-parietal and motor-planning regions rather than early visual areas. The researchers found that LAMs displayed action-dominant representations especially in frontal-motor cortex, while VLMs showed more balanced representations between action and reasoning processes.


This research bridges neuroscience and AI by revealing how modern foundation models process interactive tasks in ways that correspond to human brain function. Understanding these alignment patterns could inform the development of more human-like AI systems for decision-making and planning, and provide insights into how the brain represents actions and reasoning during complex interactive behaviors.


Understand the Science

Reinforcement learning 61 articles Explore Concept → Functional magnetic resonance imaging Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models’ internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5$times$ when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.

Source: Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay