The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective

My bibliography Save this paper

The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective

Author

Listed:

George Gui
Olivier Toubia

Registered:

Abstract

Large Language Models (LLMs) have shown impressive potential to simulate human behavior. We identify a fundamental challenge in using them to simulate experiments: when LLM-simulated subjects are blind to the experimental design (as is standard practice with human subjects), variations in treatment systematically affect unspecified variables that should remain constant, violating the unconfoundedness assumption. Using demand estimation as a context and an actual experiment as a benchmark, we show this can lead to implausible results. While confounding may in principle be addressed by controlling for covariates, this can compromise ecological validity in the context of LLM simulations: controlled covariates become artificially salient in the simulated decision process, which introduces focalism. This trade-off between unconfoundedness and ecological validity is usually absent in traditional experimental design and represents a unique challenge in LLM simulations. We formalize this challenge theoretically, showing it stems from ambiguous prompting strategies, and hence cannot be fully addressed by improving training data or by fine-tuning. Alternative approaches that unblind the experimental design to the LLM show promise. Our findings suggest that effectively leveraging LLMs for experimental simulations requires fundamentally rethinking established experimental design practices rather than simply adapting protocols developed for human subjects.

Suggested Citation

George Gui & Olivier Toubia, 2023. "The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective," Papers 2312.15524, arXiv.org, revised Jan 2025.

Handle: RePEc:arx:papers:2312.15524

Download full text from publisher

References listed on IDEAS

John J. Horton, 2023. "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?," NBER Working Papers 31122, National Bureau of Economic Research, Inc.
Günter J. Hitsch & Ali Hortaçsu & Xiliang Lin, 2021. "Prices and promotions in U.S. retail markets," Quantitative Marketing and Economics (QME), Springer, vol. 19(3), pages 289-368, December.
John J. Horton, 2023. "Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?," Papers 2301.07543, arXiv.org.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Hortense Fong & George Gui, 2024. "Modeling Story Expectations to Understand Engagement: A Generative Framework Using LLMs," Papers 2412.15239, arXiv.org, revised Mar 2025.
Ruicheng Ao & Hongyu Chen & David Simchi-Levi, 2024. "Prediction-Guided Active Experiments," Papers 2411.12036, arXiv.org, revised Nov 2024.
Ali Goli & Amandeep Singh, 2024. "Frontiers: Can Large Language Models Capture Human Preferences?," Marketing Science, INFORMS, vol. 43(4), pages 709-722, July.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Kevin Leyton-Brown & Paul Milgrom & Neil Newman & Ilya Segal, 2024. "Artificial Intelligence and Market Design: Lessons Learned from Radio Spectrum Reallocation," NBER Chapters, in: New Directions in Market Design, National Bureau of Economic Research, Inc.
Kirshner, Samuel N., 2024. "GPT and CLT: The impact of ChatGPT's level of abstraction on consumer recommendations," Journal of Retailing and Consumer Services, Elsevier, vol. 76(C).
Zengqing Wu & Run Peng & Xu Han & Shuyuan Zheng & Yixin Zhang & Chuan Xiao, 2023. "Smart Agent-Based Modeling: On the Use of Large Language Models in Computer Simulations," Papers 2311.06330, arXiv.org, revised Dec 2023.
Holtdirk, Tobias & Assenmacher, Dennis & Bleier, Arnim & Wagner, Claudia, 2024. "Fine-Tuning Large Language Models to Simulate German Voting Behaviour (Working Paper)," OSF Preprints udz28_v1, Center for Open Science.
Joshua C. Yang & Damian Dailisan & Marcin Korecki & Carina I. Hausladen & Dirk Helbing, 2024. "LLM Voting: Human Choices and AI Collective Decision Making," Papers 2402.01766, arXiv.org, revised Aug 2024.
Nir Chemaya & Daniel Martin, 2023. "Perceptions and Detection of AI Use in Manuscript Preparation for Academic Journals," Papers 2311.14720, arXiv.org, revised Jan 2024.
Lijia Ma & Xingchen Xu & Yong Tan, 2024. "Crafting Knowledge: Exploring the Creative Mechanisms of Chat-Based Search Engines," Papers 2402.19421, arXiv.org.
Ali Goli & Amandeep Singh, 2023. "Exploring the Influence of Language on Time-Reward Perceptions in Large Language Models: A Study Using GPT-3.5," Papers 2305.02531, arXiv.org, revised Jun 2023.
Cova, Joshua & Schmitz, Luuk, 2024. "A primer for the use of classifier and generative large language models in social science research," OSF Preprints r3qng_v1, Center for Open Science.
Evangelos Katsamakas, 2024. "Business models for the simulation hypothesis," Papers 2404.08991, arXiv.org.
Yuan Gao & Dokyun Lee & Gordon Burtch & Sina Fazelpour, 2024. "Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina," Papers 2410.19599, arXiv.org, revised Jan 2025.
Jiaxin Liu & Yi Yang & Kar Yan Tam, 2025. "Evaluating and Aligning Human Economic Risk Preferences in LLMs," Papers 2503.06646, arXiv.org.
Christoph Engel & Max R. P. Grossmann & Axel Ockenfels, 2023. "Integrating machine behavior into human subject experiments: A user-friendly toolkit and illustrations," Discussion Paper Series of the Max Planck Institute for Research on Collective Goods 2024_01, Max Planck Institute for Research on Collective Goods.
- Christoph Engel & Max R. P. Grossmann & Axel Ockenfels, 2024. "Integrating Machine Behavior into Human Subject Experiments: A User-Friendly Toolkit and Illustrations," ECONtribute Discussion Papers Series 302, University of Bonn and University of Cologne, Germany.
Yiting Chen & Tracy Xiao Liu & You Shan & Songfa Zhong, 2023. "The emergence of economic rationality of GPT," Proceedings of the National Academy of Sciences, Proceedings of the National Academy of Sciences, vol. 120(51), pages 2316205120-, December.
- Yiting Chen & Tracy Xiao Liu & You Shan & Songfa Zhong, 2023. "The Emergence of Economic Rationality of GPT," Papers 2305.12763, arXiv.org, revised Nov 2023.
Samuel Chang & Andrew Kennedy & Aaron Leonard & John A. List, 2024. "12 Best Practices for Leveraging Generative AI in Experimental Research," NBER Working Papers 33025, National Bureau of Economic Research, Inc.
- Samuel Chang & Andrew Kennedy & Aaron Leonard & John List, 2024. "12 Best Practices for Leveraging Generative AI in Experimental Research," Artefactual Field Experiments 00796, The Field Experiments Website.
Jiafu An & Difang Huang & Chen Lin & Mingzhu Tai, 2024. "Measuring Gender and Racial Biases in Large Language Models," Papers 2403.15281, arXiv.org.
Daniel Albert & Stephan Billinger, 2024. "Reproducing and Extending Experiments in Behavioral Strategy with Large Language Models," Papers 2410.06932, arXiv.org.
Fulin Guo, 2023. "GPT in Game Theory Experiments," Papers 2305.05516, arXiv.org, revised Dec 2023.
Iuliia Alekseenko & Dmitry Dagaev & Sofia Paklina & Petr Parshakov, 2025. "Strategizing with AI: Insights from a Beauty Contest Experiment," Papers 2502.03158, arXiv.org, revised Apr 2025.
Cova, Joshua & Schmitz, Luuk, 2024. "A primer for the use of classifier and generative large language models in social science research," OSF Preprints r3qng, Center for Open Science.

More about this item

NEP fields

This paper has been announced in the following NEP Reports:

NEP-AIN-2024-01-15 (Artificial Intelligence)
NEP-CMP-2024-01-15 (Computational Economics)
NEP-ECM-2024-01-15 (Econometrics)
NEP-EXP-2024-01-15 (Experimental Economics)

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2312.15524. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: http://arxiv.org/ .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

NEP fields

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data