Last updated: 18 September 2026
Quick Answer: Trust real respondents for the decision and synthetic users for the exploration before it. The deciding axis is your unit of analysis, because synthetic respondents track aggregate patterns far better than they track individuals. A July 2026 cross-domain benchmark found models picked the wrong segment in half of US cases.
The split is not about quality in general. It is about which unit of analysis your decision rests on. Aggregate accuracy and individual accuracy are separate measurements. A September 2026 evaluation paper from marketing academics argues that the wild range in reported synthetic accuracy, from near-perfect to near-chance, mostly reflects which of the two a vendor chose to report.
That is the finding to carry into a vendor call. Aggregate measures often perform well even when the model was given very little information, and they can mask a complete absence of respondent-level differentiation.
Four different things get sold under this label. Rows are alphabetical by approach, so the order carries no ranking.
| Approach | What grounds it | What the evidence supports | Where it breaks |
|---|---|---|---|
| Augmented panel (synthetic boost) | A real survey of 300 or more respondents | Reading a segment too thin to read on its own | Segments above 15% of the field, and questions asked only of that segment |
| Digital twin | Respondent-level data from the people being imitated | Filling a question a fielded study forgot to ask | Questions the grounding data cannot predict |
| Real respondents | People, recruited and verified | Any decision that turns on individuals or novelty | Cost, field time, and panel fraud |
| Ungrounded persona | A demographic profile in a prompt | Exploration and stimulus pretesting | Individual-level accuracy, and the spread of real opinion |
What Is a Synthetic User, and What Is It Actually Made Of?
A synthetic user is a language model answering research questions as if it were a person with a stated profile. What it is made of decides what it can be trusted for, and the September 2026 evaluation paper separates three grades of grounding: ungrounded model responses, segment-level personas, and individual-level digital twins.
Ungrounded personas are the cheapest form. The model receives a demographic profile and nothing else, so every answer comes from patterns in pre-training rather than from anyone in your category.
Segment-level personas add aggregate survey data about a group. Individual-level twins go further and condition the model on respondent-level records from the actual people being imitated.
Those three are not interchangeable, and the distinction is the first question to put to a vendor. The platform pages rarely make it for you, which is why the platforms selling synthetic respondents need reading against their own grounding claims rather than their accuracy claims.
Where Do Synthetic Respondents Match Real Data, and Where Do They Fail?
They match on direction and aggregate distribution, and they fail on individuals. A July 2026 cross-domain benchmark ran four models, spanning two families and an 8B-to-frontier capability range, against the General Social Survey and the World Values Survey, and found no model beat even the strongest non-LLM baseline at the individual level.
What the 2026 Benchmarks Measured
Two failures replicated across both domains, all four models and both model families. Models over-determine demographics, treating identity as far more predictive of attitudes than it is among real people. Neither failure was fixed by a larger, more capable model.
The decision-impact analysis is the part worth quoting in a planning meeting. On a segment-targeting task, the models inflated between-segment gaps two to fourfold. They would have directed a team to the wrong segment in half of US cases and most cross-cultural ones, and they manufactured segment splits that do not exist in real people.
Why Answer Ordering Changed the Earlier Results
There is an older result underneath this. Evaluating 43 language models against American Community Survey questions from the US Census Bureau, researchers found responses governed by ordering and labeling biases, such as a pull toward the option labeled "A". Once answer ordering was randomized to correct for that, models trended toward uniformly random responses, irrespective of model size or pre-training data.
Which Tasks Should Use Which Respondents?
Rows run in the order a study usually runs, from exploration to the final decision, so the order carries no ranking of either approach.
| Research task | Better buy in 2026 | Why |
|---|---|---|
| Early idea and concept exploration | Synthetic | Speed and volume matter more than precision at this stage |
| Questionnaire and stimulus pretesting | Synthetic | You are testing the instrument, not measuring a population |
| Reading a thin segment in an existing quant study | Augmented panel | A modeled boost is cheaper than re-fielding a rare target |
| Segmentation and targeting | Real | Models manufacture segment splits that do not exist |
| Magnitude, pricing and purchase intent | Real | Synthetic estimates overstate size and positivity |
| A new product, category or market | Real | Pre-training saw no one who has used the thing |
| Legal, regulatory or financial consequence | Real | No professional code treats generated rows as observed data |
Read that table honestly and the synthetic case is real. For screening 30 message variants down to five, or checking whether a question reads the way you meant it to, a synthetic panel is the better buy and a human sample is a waste of fieldwork budget. Save the fieldwork for the shortlist and for what purchase intent scores actually predict, which is where overstated magnitude does real damage.
How Do You Calibrate a Synthetic Panel Against Real Interviews?
You calibrate by holding back real data and scoring the model against it. The September 2026 paper adds a cheaper screen that needs no ground truth at all. Run a random forest predicting the twin's outputs from the data used to build the twin, then read the R-squared as an answerability score for each question.
Across 108 attitude questions from a nationally representative survey of 3,063 people, screening at R-squared above 0.7 raised the mean twin-to-human individual correlation by 15% and cut the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment screened in the same direction but more weakly.
The practical sequence is short.
- Ground the model on your own respondent-level data.
- Hold out a real sample the model never saw.
- Score question by question, not study by study.
- Field real interviews for everything that fails the screen. That last step is why where good survey respondents come from decides the ceiling on a synthetic program rather than sitting beside it.
Speed is what makes that loop affordable. Alchemic runs AI-moderated interviews as text natively inside WhatsApp, with voice notes supported and no link or app, and as voice and video interviews on the browser. A 200-interview qualitative study turns around in about three days, which puts a real validation sample inside the same sprint as the synthetic read. The reliability conditions are covered separately in when AI-moderated interviews produce reliable data.
What Do the Standards Bodies Say About Synthetic Respondents?
They say synthetic samples work inside stated boundaries and break outside them. An ESOMAR Congress 2024 paper by Samuel Cohen and Thomas Duhard set out thresholds from more than 7,000 parallel tests across 40 Pew American Trends Panel datasets, covering 7,316 segments with a mean size of 48.
The two hard limits are worth writing into a brief. The model must be trained on an original field of at least 300 respondents, and the boosted segment should sit between 1% and 15% of the total field.
Within those limits the gains are measurable. The mean effective sample size was 2.855, meaning a boost was statistically worth roughly 2.9 times the real data it started from, running up to 3.5 times on smaller segments and about 2.5 times on larger ones.
The same paper names the statistical catch. The assumptions behind t-tests and similar tests do not naturally apply to synthetic samples, so significance testing has to be rebuilt with bootstrapped confidence intervals or handled conservatively. Disclosure is now a separate obligation, covered in consent and disclosure rules for AI-moderated research.
How Do Cost and Speed Actually Compare in 2026?
Synthetic wins on both, and no independent primary source publishes a like-for-like price benchmark, so treat every published ratio as a vendor figure. What is documented is the cost structure on the human side and the specific jobs a boost removes.
Probability-based recruitment is expensive because it is time-intensive and labor-intensive. Pew Research Center recruits offline by mail, and its own account of the trade-off is that rigor is what costs money, not the interview itself.
A synthetic boost removes a different cost: the incremental field. In the ESOMAR case study, a French election survey of 8,000 adults reported the statistical equivalent of 580 secondary school teachers derived from 116 real interviews, a target that would otherwise have needed its own fieldwork. Broader budget shapes sit in market research costs and pricing models.
Speed on the human side has moved too, which narrows the gap that made synthetic panels attractive in the first place. Both sides of that comparison are laid out in AI versus human moderated interviews and in the wider sweep of AI market research tools for consumer insights.
What Neither Approach Can Settle
Neither approach settles whether your respondents are human, which sounds like a synthetic problem and is not. A 2025 paper in the Proceedings of the National Academy of Sciences built an autonomous agent that passed 99.8% of 6,000 standard attention-check trials while posing as a human respondent in online survey research.
Common Mistakes to Avoid
- Reporting a synthetic estimate without labeling it. Professional codes require the provenance to be stated.
- Validating on aggregates alone. Aggregate accuracy hides a complete absence of individual differentiation.
- Trusting a bigger model to fix it. Neither failure mode was solved by more capability.
- Using generated rows for a consequential claim. No code treats them as observed data.
The same agent could be instructed to skew a poll, and more quietly, could infer a researcher's latent hypothesis and produce data that confirmed it. A low-barrier human panel with that problem has the same validity gap as a synthetic one.
Neither approach settles novelty either. A model cannot report on a product nobody has used, and a panel recruited from people who have never encountered your category will not either. That is a recruitment question, covered in sample validity and who a study misses.
Reach is the third thing neither method fixes on its own. Alchemic fields in 14 markets including the USA and the UK, publishes 57+ languages including Spanish, Arabic and Mandarin, and runs managed fieldwork or bring your own, which changes who is reachable rather than how well anyone is simulated. Early-stage screening remains the synthetic strength, and concept testing with real shoppers or full AI-moderated qualitative interviews at scale remain the place a shortlist gets decided.

