AI & Computational Science

Can AI Chatbots Accurately Mimic Human Survey Responses?

How the science connects

Large language modelSurvey methodology

AI Insight

Researchers compared survey responses from 420 human software developers in Silicon Valley with synthetic responses generated by five leading AI language models attempting to replicate the same survey. While the AI models produced technically plausible and consistent results that aligned with each other, they failed to capture counterintuitive insights present in human responses and instead reproduced conventional wisdom. The study found that all AI models deviated from human data in similar ways, making the actual human responses appear as outliers.


This research challenges the growing practice of using AI-generated synthetic data to replace or supplement human survey responses in organizational and social research. The findings suggest synthetic surveys may be useful for identifying societal assumptions and conventional wisdom but cannot substitute for actual human participants when seeking novel insights about social beliefs, particularly in understudied contexts.


Understand the Science

Large language model 68 articles Explore Concept → Survey methodology Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

-cross
Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.

Source: Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data