Home Feeds Careers Get in Touch

Synthetic Users vs Real Respondents in Research 2026

synthetic users synthetic respondents ai focus groups ai market research
Banner: Synthetic users vs real respondents and sample validity in consumer research

TL;DR

  • Synthetic respondents reproduce aggregate patterns reasonably well and fail on the individuals, segments and edge cases a decision usually turns on.
  • A July 2026 cross-domain benchmark found no language model beat the strongest non-LLM baseline at the individual level, and models sent teams to the wrong segment in half of US cases.
  • Use synthetic panels to explore and pretest cheaply.
  • Use real respondents to decide, to ground the model, and to check what it produced.

Last updated: 18 September 2026

Quick Answer: Trust real respondents for the decision and synthetic users for the exploration before it. The deciding axis is your unit of analysis, because synthetic respondents track aggregate patterns far better than they track individuals. A July 2026 cross-domain benchmark found models picked the wrong segment in half of US cases.

The split is not about quality in general. It is about which unit of analysis your decision rests on. Aggregate accuracy and individual accuracy are separate measurements. A September 2026 evaluation paper from marketing academics argues that the wild range in reported synthetic accuracy, from near-perfect to near-chance, mostly reflects which of the two a vendor chose to report.

That is the finding to carry into a vendor call. Aggregate measures often perform well even when the model was given very little information, and they can mask a complete absence of respondent-level differentiation.

Four different things get sold under this label. Rows are alphabetical by approach, so the order carries no ranking.

Approach What grounds it What the evidence supports Where it breaks
Augmented panel (synthetic boost) A real survey of 300 or more respondents Reading a segment too thin to read on its own Segments above 15% of the field, and questions asked only of that segment
Digital twin Respondent-level data from the people being imitated Filling a question a fielded study forgot to ask Questions the grounding data cannot predict
Real respondents People, recruited and verified Any decision that turns on individuals or novelty Cost, field time, and panel fraud
Ungrounded persona A demographic profile in a prompt Exploration and stimulus pretesting Individual-level accuracy, and the spread of real opinion

What Is a Synthetic User, and What Is It Actually Made Of?

A synthetic user is a language model answering research questions as if it were a person with a stated profile. What it is made of decides what it can be trusted for, and the September 2026 evaluation paper separates three grades of grounding: ungrounded model responses, segment-level personas, and individual-level digital twins.

Ungrounded personas are the cheapest form. The model receives a demographic profile and nothing else, so every answer comes from patterns in pre-training rather than from anyone in your category.

Segment-level personas add aggregate survey data about a group. Individual-level twins go further and condition the model on respondent-level records from the actual people being imitated.

Those three are not interchangeable, and the distinction is the first question to put to a vendor. The platform pages rarely make it for you, which is why the platforms selling synthetic respondents need reading against their own grounding claims rather than their accuracy claims.

Where Do Synthetic Respondents Match Real Data, and Where Do They Fail?

They match on direction and aggregate distribution, and they fail on individuals. A July 2026 cross-domain benchmark ran four models, spanning two families and an 8B-to-frontier capability range, against the General Social Survey and the World Values Survey, and found no model beat even the strongest non-LLM baseline at the individual level.

What the 2026 Benchmarks Measured

Two failures replicated across both domains, all four models and both model families. Models over-determine demographics, treating identity as far more predictive of attitudes than it is among real people. Neither failure was fixed by a larger, more capable model.

The decision-impact analysis is the part worth quoting in a planning meeting. On a segment-targeting task, the models inflated between-segment gaps two to fourfold. They would have directed a team to the wrong segment in half of US cases and most cross-cultural ones, and they manufactured segment splits that do not exist in real people.

Why Answer Ordering Changed the Earlier Results

There is an older result underneath this. Evaluating 43 language models against American Community Survey questions from the US Census Bureau, researchers found responses governed by ordering and labeling biases, such as a pull toward the option labeled "A". Once answer ordering was randomized to correct for that, models trended toward uniformly random responses, irrespective of model size or pre-training data.

Which Tasks Should Use Which Respondents?

Rows run in the order a study usually runs, from exploration to the final decision, so the order carries no ranking of either approach.

Research task Better buy in 2026 Why
Early idea and concept exploration Synthetic Speed and volume matter more than precision at this stage
Questionnaire and stimulus pretesting Synthetic You are testing the instrument, not measuring a population
Reading a thin segment in an existing quant study Augmented panel A modeled boost is cheaper than re-fielding a rare target
Segmentation and targeting Real Models manufacture segment splits that do not exist
Magnitude, pricing and purchase intent Real Synthetic estimates overstate size and positivity
A new product, category or market Real Pre-training saw no one who has used the thing
Legal, regulatory or financial consequence Real No professional code treats generated rows as observed data

Read that table honestly and the synthetic case is real. For screening 30 message variants down to five, or checking whether a question reads the way you meant it to, a synthetic panel is the better buy and a human sample is a waste of fieldwork budget. Save the fieldwork for the shortlist and for what purchase intent scores actually predict, which is where overstated magnitude does real damage.

How Do You Calibrate a Synthetic Panel Against Real Interviews?

You calibrate by holding back real data and scoring the model against it. The September 2026 paper adds a cheaper screen that needs no ground truth at all. Run a random forest predicting the twin's outputs from the data used to build the twin, then read the R-squared as an answerability score for each question.

Across 108 attitude questions from a nationally representative survey of 3,063 people, screening at R-squared above 0.7 raised the mean twin-to-human individual correlation by 15% and cut the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment screened in the same direction but more weakly.

The practical sequence is short.

  1. Ground the model on your own respondent-level data.
  2. Hold out a real sample the model never saw.
  3. Score question by question, not study by study.
  4. Field real interviews for everything that fails the screen. That last step is why where good survey respondents come from decides the ceiling on a synthetic program rather than sitting beside it.

Speed is what makes that loop affordable. Alchemic runs AI-moderated interviews as text natively inside WhatsApp, with voice notes supported and no link or app, and as voice and video interviews on the browser. A 200-interview qualitative study turns around in about three days, which puts a real validation sample inside the same sprint as the synthetic read. The reliability conditions are covered separately in when AI-moderated interviews produce reliable data.

What Do the Standards Bodies Say About Synthetic Respondents?

They say synthetic samples work inside stated boundaries and break outside them. An ESOMAR Congress 2024 paper by Samuel Cohen and Thomas Duhard set out thresholds from more than 7,000 parallel tests across 40 Pew American Trends Panel datasets, covering 7,316 segments with a mean size of 48.

The two hard limits are worth writing into a brief. The model must be trained on an original field of at least 300 respondents, and the boosted segment should sit between 1% and 15% of the total field.

Within those limits the gains are measurable. The mean effective sample size was 2.855, meaning a boost was statistically worth roughly 2.9 times the real data it started from, running up to 3.5 times on smaller segments and about 2.5 times on larger ones.

The same paper names the statistical catch. The assumptions behind t-tests and similar tests do not naturally apply to synthetic samples, so significance testing has to be rebuilt with bootstrapped confidence intervals or handled conservatively. Disclosure is now a separate obligation, covered in consent and disclosure rules for AI-moderated research.

How Do Cost and Speed Actually Compare in 2026?

Synthetic wins on both, and no independent primary source publishes a like-for-like price benchmark, so treat every published ratio as a vendor figure. What is documented is the cost structure on the human side and the specific jobs a boost removes.

Probability-based recruitment is expensive because it is time-intensive and labor-intensive. Pew Research Center recruits offline by mail, and its own account of the trade-off is that rigor is what costs money, not the interview itself.

A synthetic boost removes a different cost: the incremental field. In the ESOMAR case study, a French election survey of 8,000 adults reported the statistical equivalent of 580 secondary school teachers derived from 116 real interviews, a target that would otherwise have needed its own fieldwork. Broader budget shapes sit in market research costs and pricing models.

Speed on the human side has moved too, which narrows the gap that made synthetic panels attractive in the first place. Both sides of that comparison are laid out in AI versus human moderated interviews and in the wider sweep of AI market research tools for consumer insights.

What Neither Approach Can Settle

Neither approach settles whether your respondents are human, which sounds like a synthetic problem and is not. A 2025 paper in the Proceedings of the National Academy of Sciences built an autonomous agent that passed 99.8% of 6,000 standard attention-check trials while posing as a human respondent in online survey research.

Common Mistakes to Avoid

  • Reporting a synthetic estimate without labeling it. Professional codes require the provenance to be stated.
  • Validating on aggregates alone. Aggregate accuracy hides a complete absence of individual differentiation.
  • Trusting a bigger model to fix it. Neither failure mode was solved by more capability.
  • Using generated rows for a consequential claim. No code treats them as observed data.

The same agent could be instructed to skew a poll, and more quietly, could infer a researcher's latent hypothesis and produce data that confirmed it. A low-barrier human panel with that problem has the same validity gap as a synthetic one.

Neither approach settles novelty either. A model cannot report on a product nobody has used, and a panel recruited from people who have never encountered your category will not either. That is a recruitment question, covered in sample validity and who a study misses.

Reach is the third thing neither method fixes on its own. Alchemic fields in 14 markets including the USA and the UK, publishes 57+ languages including Spanish, Arabic and Mandarin, and runs managed fieldwork or bring your own, which changes who is reachable rather than how well anyone is simulated. Early-stage screening remains the synthetic strength, and concept testing with real shoppers or full AI-moderated qualitative interviews at scale remain the place a shortlist gets decided.

Frequently Asked Questions

What is a synthetic user?
A synthetic user is a language model producing research answers under an assigned profile. The label is not standardized across fields. In product and UX work it usually means a persona interviewed in conversation. In quantitative research it usually means a modeled survey row added to a dataset. Those two things are validated differently, so ask which one a vendor means before comparing any accuracy figure.
Do synthetic respondents need consent, and what must you disclose?
No, because a generated profile is not a person. It gives no consent, takes no incentive, and cannot withdraw, so human-subjects protections do not attach to it. The obligations shift instead to the report: professional codes now require research outputs to state which findings came from generated profiles and which came from people who answered.
Can a model predict an opinion nobody has ever been asked about?
Only weakly, because this kind of inference works better backward than forward. A study using General Social Survey data from 1972 to 2021 found language models retrodicted masked historical opinions strongly, while performance on entirely unasked opinions remained modest. Filling a known gap is a different task from predicting something nobody has ever been asked.
What are synthetic focus groups and how do they work in market research?
Several model agents are given personas and prompted to discuss a stimulus together. They produce transcripts quickly and miss what a group is for. Nobody changes position under social pressure, and studies of AI-generated opinion find it stereotypes groups and understates the level of disagreement in a real population. Treat the output as a hypothesis list, not as a reading of consensus.
What are the limits of boosting a thin segment with modeled rows?
Two limits are documented and easy to miss, and both are narrower than the technique sounds. A boost cannot answer questions that were only asked of the target segment, because the model has no other respondents to learn from. And synthetic rows behave like the people interviewed on the original field date, so events since that date are absent.
Is survey data accurate?
It depends on recruitment, and both failure modes are measurable. In one Pew Research Center experiment, 12% of opt-in respondents under 30 claimed a license to operate a nuclear submarine, against a real share that rounds to zero. Automated respondents leave a different trace, sometimes answering general-knowledge questions more accurately than any human population would.
How much does survey research cost?
Incentives are a small share of it. Panelists on one large probability-based US panel take fewer than two surveys a month and receive an average of $11 per survey, so the money sits in recruitment, weighting and field management rather than in respondent payments. That is also why a rare target costs far more per completed interview than a general population sample.

About the Author

Sreenadh Narayanan is the founder of Alchemic, an AI-powered consumer research platform used for ad testing, concept testing and brand tracking. He writes Alchemic's guides on qualitative research and research methods, covering interview design, sample sizes and how teams turn customer conversations into decisions.