The cross-domain benchmark and the evaluation framework for intelligent synthetic-user evidence is made available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.
Abstract
Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.
Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as's...
Indira Sen, Georg Ahnert, Leah von der Heyde et al.· 0 citations
Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.
Yuanzi Li, Jun-Hao Wang, Minghui Liu et al.· 0 citations
A definition of LLM-based simulation of human samples is offered and the applications of this method across four major domains—psychometrics and machine psychology, consumer and market research, experimental simulation and digital twin construction, and the simulation of political opinion and social sentiment are summa...
Yuying Zhang· Journal of Artificial Intell...· 0 citations
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but als...
Yong Peng, Qing-Shui Gu, Li-Ya Zhu et al.· 0 citations
Global stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.
Evaluating six matched base and post-trained models on the Pew American Trends Panel, it is found that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure.
Seth Grief-Albert, Jessica Y. Bo, Di-Fan Jiao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.