AI Insight
Researchers have developed BehaviorBench, a comprehensive benchmark system to evaluate how well AI foundation models perform on behavioral science tasks across psychology, sociology, and economics. The evaluation assesses models on four capabilities: behavior prediction, strategic decision-making, trait inference, and behavioral knowledge application, measuring both individual-level accuracy and population-level alignment. Testing revealed that current leading AI models struggle with these tasks, with general-purpose language models underestimating human response diversity while specialized behavior models perform poorly on individual predictions, though fine-tuning on behavioral data can improve both metrics.
Why it matters
This benchmark provides the first systematic framework for assessing AI models intended for behavioral science applications, which is crucial as these models are increasingly used to simulate human subjects and predict behavior in research and commercial contexts. The finding that individual and population-level performance diverge highlights potential risks in deploying AI for behavioral predictions without proper evaluation.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject accuracy but also population-level alignment, an essential requirement for behavioral validity. Our evaluation shows that BehaviorBench remains challenging for leading general-purpose LLMs and behavior foundation models that are specifically trained with behavioral data. We find that individual-level and distributional performance do not always align. General-purpose LLMs tend to underestimate the diversity of human responses, whereas behavior foundation models often lag behind at individual-level prediction. Our investigation further demonstrates how fine-tuning on diverse behavioral data can improve both individual-level prediction and distributional alignment, balancing these two objectives. Our results highlight the importance of evaluation at both individual and distributional levels, establishing BehaviorBench as a foundation for developing and assessing behaviorally aligned AI systems. Our BehaviorBench and models can be accessed via https://umich-foreseer.github.io/behaviorbench/
Source: BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks