Lode

Stand on the shoulders of giants.

Open the curator →
Source
arXiv
Published
Runtime
0:00
Snippets
7

A conversation between

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Waveform of the source interview with highlighted segments per snippet.
0:00 0:00

§02

Snippets

  1. LLMs reproduce qualitative trends but fail to preserve the joint distributions, reliability, and mediation pathways of real survey data—a Gaussian-copula baseline outperforms all 37 tested models on psychometric structure.

    Survey datasets drive organizational decisions and social-science research; synthetic data that passes surface plausibility can introduce silent bias downstream.

  2. Counterfactual demographic swaps reveal education-driven effects (d=0.56) that dwarf gender (0.12) and role (0.18), showing LLMs amplify some demographic signals while dampening others.

    If synthetic respondents misweight demographic effects, studies built on them will misestimate who is most affected by organizational interventions.

  3. The LLM 'crowd' shows mean inter-model similarity of 0.73, higher than similarity to humans, and models are more similar to each other than to any individual human respondent.

    Synthetic crowds collapse diversity—different models converge on the same wrong answers, defeating the value of a large sample.

  4. Regressors trained on synthetic LLM responses lose predictive validity on held-out humans (mean R² −0.18 vs. 0.28 for human-trained models).

    Synthetic data can flip a valid predictor into a negative one—a warning for downstream applied modeling.

  5. Every LLM exhibits a strong acquiescence shift (+0.84 SD), a systematic bias toward agreement that dwarfs human response patterns.

    This uniform distortion means synthetic data will overestimate the strength of agreement and miss disagreement, skewing substantive conclusions.

  6. LLMs fabricate indirect effects on 3 of 10 placebo mediation paths—falsely inferring causal chains that don't exist in the data.

    If synthetic data invents false mechanisms, researchers may pursue interventions based on nonexistent psychological pathways.

  7. Persona-disclosure intensity, presentation format, and reasoning-effort ablations show LLMs do not converge toward human response patterns as conditioning increases.

    Simply giving models more context does not solve the fundamental mismatch—the problem is structural, not superficial.

§03

Synthesis

LLMs Sound Convincing as Survey Respondents—But Their Data Doesn't Work

Individual LLM responses to survey questions look plausible. The problem lies deeper: when researchers use LLMs to generate synthetic survey data, the resulting dataset fails on every measure that matters to psychologists—reliability, latent structure, demographic realism, and predictive validity. This paper audits 37 LLM models against real human survey data and finds that plausibility masks fundamental misalignment with how human psychology actually works.

The Test

The authors collected Lithuanian organisational-psychology survey responses from 263 real employees (three validated instruments, 68 items total) and then asked LLMs to role-play as synthetic respondents. They conditioned models on realistic profiles and varied how much information was disclosed about respondent background. The key question: does the distribution of LLM responses match the human distribution, or just the surface appearance?

They measured this using a Psychometric Similarity Score (PSS) anchored against non-LLM baselines and a human-versus-human ceiling. The test suite was ruthless: it checked whether correlations between survey subscales matched, whether latent factor structures aligned, whether counterfactual edits (swapping gender or education) produced realistic effect sizes, and whether predictive models trained on synthetic data worked on real humans.

What They Found

LLMs reproduce the direction of human relationships—they get the big picture roughly right. But a simple Gaussian-copula statistical baseline beats every LLM model on sample-driven similarity. More damning: LLMs are far more similar to each other (mean inter-LLM similarity 0.73) than to humans, suggesting they converge on a shared mode rather than sampling the human distribution.

Counterfactual demographic swaps revealed a structural problem. Education effects were enormous (effect size 0.56), while gender and role effects were tiny (0.12 and 0.18 respectively)—the reverse of human patterns. All 37 models exhibited a strong acquiescence bias: they agree with survey statements 0.84 standard deviations more than humans do. When researchers trained predictive regression models on the synthetic data and tested them on real humans, accuracy collapsed (R² dropped from 0.28 to −0.18). In 3 of 10 mediation analyses, LLM samples fabricated indirect effects that don't exist in real data.

Memorization played no role—models with higher verbatim recall of training data showed zero correlation with leaderboard ranking.

Why It Matters

Using LLMs as a cheap substitute for collecting human survey data can introduce subtle but serious bias into downstream research. A study trained on synthetic respondents will report false correlations, overestimate certain demographic effects, and produce predictive models that fail when applied to actual people. These errors are harder to catch than obviously wrong answers because individual responses read naturally and pass surface-level plausibility checks.

The authors' conclusion is clear: LLM samples are not interchangeable with human data. Researchers tempted by speed and cost savings should benchmark carefully against the psychometric properties their analysis actually depends on—not just whether individual responses seem reasonable.

Mine your own.

Lode is a workbench, not a feed. Paste a YouTube URL. The model proposes a transcript, a set of quote-grounded snippets, a synthesis essay, and the fan-out. You decide what stays.

Open the curator