Briefing / 2026-09-08 / 8 min read

Synthetic Respondents vs Real Respondents: What the Vendors' Own Research Shows

A 2026 ESOMAR paper by a synthetic data vendor and Google concludes qualitative research stays with real humans. Where synthetic works and where it fails.

Black and white geometric architecture, Synthetic Respondents vs Real Respondents: What the Vendors' Own Research Shows

Synthetic respondents are AI-generated answers that stand in for real survey participants. According to a February 2026 ESOMAR paper co-authored by the synthetic data vendor Fairgen and Google, they are valid only for boosting close-ended quantitative data on top of a real sample of at least 300 people, are not appropriate as a substitute for statistical inference, and are poor at surfacing novel findings. Qualitative and open-ended research remains the domain of real human research.

Three different things called "synthetic"

The phrase covers three distinct techniques, and conflating them is where most confusion starts.

Synthetic augmentation takes a real survey and generates additional records to boost small subgroups, so a base of 40 respondents in a segment can be analysed as if it were 200. Digital twins build a model of a known population from its historical data and query the model instead of the people. LLM-based synthetic respondents ask a large language model to answer a questionnaire as if it were a member of the target audience.

What the vendors' own research concludes

In February 2026 Fairgen, a synthetic data company, and Google published a paper through ESOMAR evaluating these approaches. Because the authors sell the technology, the limits they concede carry weight.

ApproachWhere it holdsWhere it fails
AugmentationClose-ended quantitative data; real base of n at least 300Open-ended, qualitative, small or novel populations
Digital twinsStable, well-documented populationsReproduces every bias in the training data; blind to change
LLM respondentsDirectional exploration, questionnaire testingStatistical inference, novel findings, anything high-stakes

The paper's language on LLM respondents is direct: they are "not appropriate as a substitute for statistical inference". On qualitative work it is more direct still: open-ended and qualitative research "remain firmly within the domain of real human research". Primary fieldwork remains the gold standard.

The novelty problem

The most important limitation is not statistical. A model trained on the past can only recombine what it has seen. If your question is "what do buyers think of a product category that existed last year", a synthetic answer will be plausible. If your question is "what has changed since the regulation passed in March", or "why did our win rate fall last quarter", the model has no way to know, and it will produce a confident answer anyway. Decisions that matter are almost always about what is new.

Why synthetic is winning anyway

Synthetic data is popular because real online data has become unreliable. Kantar estimates that 30 to 40 percent of online survey responses are compromised by bots, fraud or professional respondents. Around 3 percent of devices complete 19 percent of all online surveys. Research published in PNAS found LLM-powered bots pass standard survey quality checks 99.8 percent of the time. Faced with panels a third full of noise, a synthetic panel can look like an improvement. It is a substitution of one unverifiable source for another.

When to use which

  • Use synthetic augmentation when you have a clean, verified quantitative sample of at least 300 and need to read small subgroups. Disclose it in the report.
  • Use LLM respondents to test a questionnaire or generate hypotheses before fieldwork. Never report their output as a finding.
  • Use real, verified humans for anything qualitative, anything B2B or expert, anything novel, and anything where the decision is expensive to reverse.

How Theory works

Theory Intelligence uses only verified human sources for evidence. Every interviewee is identity-checked and role-confirmed, every transcript is logged, and the client receives a provenance appendix. We use AI for scheduling, transcription and first-pass coding, and we say so. See how a Decision Study compares to synthetic research, or read The Theory Standard.

Sources

Cite thisTheory Intelligence, "Synthetic Respondents vs Real Respondents: What the Vendors' Own Research Shows", theoryintelligence.com, 2026-09-08. https://theoryintelligence.com/briefing/synthetic-respondents-vs-real-respondents/
FAQ

Questions this article answers

Are synthetic respondents accurate?

For close-ended quantitative augmentation on a real base of 300 or more, they can be. For statistical inference, qualitative research or novel questions, the vendors' own ESOMAR research says they are not appropriate.

Can synthetic respondents replace focus groups or interviews?

No. The February 2026 Fairgen and Google paper states that qualitative and open-ended research remain within the domain of real human research.

What is the difference between synthetic augmentation and synthetic respondents?

Augmentation adds generated records to a real survey to boost small subgroups. Synthetic respondents replace the survey entirely with model-generated answers. The first has a validated use; the second is directional at best.

Have a decision this applies to?

A Signal Brief answers one sharp question with verified evidence in ten business days, for a fixed fee.

Start a brief