Skip to content
Review Open access

Silicon Samples: A Review and Outlook on Large Language Models Simulating Human Respondents

2026 · Journal of Artificial Intelligence Practice · 0 citations · 22 references

TL;DR

A definition of LLM-based simulation of human samples is offered and the applications of this method across four major domains—psychometrics and machine psychology, consumer and market research, experimental simulation and digital twin construction, and the simulation of political opinion and social sentiment are summarized.

Abstract

: The use of large language models (LLMs) to simulate human respondents (silicon samples), as an emerging research topic, is attracting growing scholarly attention. By reviewing 25 representative studies from both domestic and international literature, this paper offers a definition of LLM-based simulation of human samples and examines its foundations, characteristics, and purposes. It summarizes the applications of this method across four major domains—psychometrics and machine psychology, consumer and market research, experimental simulation and digital twin construction, and the simulation of political opinion and social sentiment. It further reviews the existing evidence regarding the method's validity, along with the attendant debates, along three dimensions: psychometric validity, group fidelity and context dependence, and ethical and epistemic justice risks. Finally, the paper synthesizes a research framework for LLM-based simulation of human samples and discusses future research directions, with the aim of providing a reference for research and applied practice in this field.

Read PDF

Similar papers

Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
Review Sep 2026

Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models

Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as's...

Indira Sen, Georg Ahnert, Leah von der Heyde et al. · 0 citations
Open access Aug 2026

RoCulturaMCQ: Building a Benchmark While Learning Statistics

A pilot project in which students in a statistics course within a data science engineering program created culturally diverse multiple-choice questions, generated answers using LLMs, and applied statistical methods to assess model accuracy is presented, supporting a cultural injection hypothesis.

Denis Iorga, Razvan Muntean, Mihai Masala et al. · 0 citations
Review Jul 2026

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

The cross-domain benchmark and the evaluation framework for intelligent synthetic-user evidence is made available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.

Zihang Chen, Di Zhu, L. Zheng · 2 citations
Review Sep 2026

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique the...

YunHong Yang, Mike Thelwall, Guo-Xiu He · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.