Home Feeds Careers Get in Touch

Synthetic Respondent Platforms for Consumer Research 2026

synthetic respondent platforms synthetic respondents market research synthetic users research tools ai simulated respondents digital twins consumer research synthetic survey respondents
Synthetic respondent platforms for consumer research compared on approach, grounding data and published validation

TL;DR

  • Synthetic Users, Fairgen, Evidenza, Aaru, Artificial Societies, Simile and GWI sell three different things under one label: prompted personas, data-grounded twins, and synthetic boosts to a human panel.
  • Across 285 published comparisons, synthetic and human samples agreed 24.9 percent of the time; NIM's twins matched real choices 79 percent of the time while overrating brands and compressing variance.
  • Both professional codes require validation and disclosure, which is where real respondents sit behind any synthetic study.

Last updated: 9 September 2026

Quick Answer: Synthetic respondent platforms for consumer research in 2026 include Synthetic Users, Fairgen, Evidenza, Aaru, Artificial Societies, Simile and GWI's synthetic audiences. They sell three different things: pure AI personas, AI twins grounded in real data, or synthetic boosts to a human panel. Across 285 published comparisons, synthetic and human samples agreed only 24.9 percent of the time.

A synthetic respondent is a large language model answering a survey or an interview as if it were a person with a stated profile. Its pitch is a panel that never drops out and scales in minutes. Its evidence is that it captures broad patterns, flatters well-known brands, compresses the diversity of real opinion, and in most published comparisons diverges from the humans it is meant to stand in for.

That does not make the category useless. Far from it. It makes the buying question specific: which of three different products is on offer, what real data grounds it, and which decisions it is fit for.

Which Synthetic Respondent Platforms Exist in 2026?

Seven platforms document synthetic respondents for consumer research on their own sites: Synthetic Users, Fairgen, Evidenza, Aaru, Artificial Societies, Simile and GWI. Several survey platforms, including Qualtrics, publish synthetic data features alongside their panels. Alchemic is not a synthetic respondent vendor; it is on this page as the human layer a synthetic study needs behind it.

The named platforms differ in what they are built from:

  • Persona simulation. Synthetic Users runs a multi-agent architecture in which AI participants carry personality profiles based on the OCEAN model and keep context across an interview. It positions itself as a discovery co-pilot rather than a replacement for organic research.
  • Synthetic boosts to a real panel. Fairgen's Boost generates synthetic respondents to enlarge under-represented segments of an existing survey, and its platform turns a brand's own quantitative and qualitative research into simulated respondents the brand owns.
  • Custom synthetic samples on audience data. Evidenza builds synthetic personas trained on specific audience data and delivers managed studies from them. Aaru simulates behavior at population scale to test a product, price or message before commitment. Artificial Societies simulates audiences as interacting personas.
  • Twins built on a survey base. GWI's synthetic audiences are built on one million real respondents across 54 markets, so a persona query draws on the survey behind it.
  • Human respondents, verified. Alchemic fields AI-moderated interviews with real people on WhatsApp, web or phone, and its MCP page states the position plainly: real interviews, not synthetic answers.

Pure AI Respondents, AI Twins or an Augmented Panel: Which Is Being Sold?

Three different products share the label. A pure AI respondent is a model prompted with a demographic profile and nothing else.

An AI twin is a model conditioned on real data from the people it imitates, from a survey base to a brand's own interviews. An augmented panel is a real sample with synthetic rows added to thin segments. The grounding data decides which one a vendor is selling, and the validation evidence differs for each.

Pure AI respondents

Prompted personas are the cheapest option and the least reliable. A February 2026 arXiv study tested persona prompting against more than 70,000 respondent-item pairs from the World Values Survey.

Demographic conditioning did not improve alignment on aggregate and in many cases significantly degraded it, with the largest distortions falling on underrepresented subgroups. The persona redistributes error. It does not remove it.

AI twins

Twins grounded in first-party data perform better on structured tasks. Bain reported in May 2026 that digital twins built from a consumer technology company's historical respondent-level data replicated about 90 percent of key outcomes from a prior 1,500-person conjoint study, including the most influential features and portfolio-level launch decisions. Bain's own caveat is that twins should be built on first-party data rather than a vendor's third-party data, and that the models still lack empathy.

Augmented panels

A synthetic boost adds modeled rows to a real survey to read a segment that was too small to read on its own. The 2025 revision of the ICC/ESOMAR International Code added a definition of synthetic data and requires that its use be disclosed.

The Insights Association's guidance on synthetic data, published in 2026, goes further. Synthetic data should not be presented as equivalent to observed human responses without validation, and reports must state which findings rest on which.

How Do the Named Synthetic Platforms Compare?

The table reads each platform's own site as of 9 September 2026. "Not stated" means the site does not document it. Rows are alphabetical.

Platform Approach sold Grounding data Validation against humans published Fits best Who runs it
Aaru Population-scale behavior simulation Not stated Not stated Testing a product, price or message before commitment Platform
Alchemic Human respondents, AI-moderated Managed recruitment or the brand's own list; a compounding knowledge base Not applicable; the interviews are the ground truth Grounding a twin, validating a synthetic read, any decision that needs a real person Managed research team
Artificial Societies Simulated societies of interacting personas Not stated Not stated Audience reaction to content and campaigns Platform
Evidenza Custom synthetic samples for qual and quant Specific audience data supplied per study Not stated on the home page Managed studies on hard-to-reach audiences Project-based, managed
Fairgen Synthetic boost of a real survey; simulated respondents from uploaded research The brand's own survey and qualitative data Site cites independent validations of boosts Reading thin segments in an existing quant study Self-serve platform
GWI synthetic audiences Twins on a survey base One million real respondents, 54 markets Draws directly on the survey data Persona queries and simulated focus groups on tracked markets Self-serve inside GWI
Synthetic Users Persona simulation, multi-agent, OCEAN profiles Model plus study setup; no respondent base claimed Site cites independent comparison studies at 85 to 92 percent parity and user testimonials Problem exploration, early concept and messaging screens Self-serve

Where a synthetic platform is the better choice is worth stating plainly. Screening 20 to 100 message or concept variants down to a shortlist in an afternoon is a job no human panel does at that speed, and Synthetic Users, Aaru and Evidenza are built for it. The human study belongs on the shortlist that survives. Not before.

Common Mistakes When Reading a Synthetic Respondent Page

  • Reading "parity" without the base. A parity figure means nothing until the vendor says which humans, which questions and which metric.
  • Treating a persona as a twin. A profile in a prompt is not grounding data. Ask what real responses the model was conditioned on.
  • Skipping the disclosure line. Both professional codes now require the report to say which findings came from humans and which from synthetic respondents.

What Does the Research Say About Whether Synthetic Answers Hold Up?

The research says synthetic answers reproduce direction and rank order reasonably well, overstate magnitude and positivity, and lose the spread of real opinion. Three independent bodies of evidence published between 2025 and 2026 agree on that shape. None of them is a vendor.

The Nuremberg Institute for Market Decisions' guidelines for silicon samples, by Marko Sarstedt, Susanne Adler, Lea Rau and Bernd Schmitt, report a literature review of 285 silicon-to-human comparisons. In 24.9 percent the results were similar, in 65.3 percent they diverged, and in 9.8 percent they aligned only partially. The authors' recommended uses are pretesting stimuli and survey items, extending existing quantitative samples, and synthetic personas for exploratory qualitative work.

NIM's own digital twin experiment, by Carolin Kaiser and colleagues, ran a brand funnel with synthetic respondents against real ones. The twins matched real choices about 79 percent of the time and reproduced the overall pattern. They also overestimated selection of well-known brands, rated brands more positively by an average of 1.2 points on a seven-point scale, and showed significantly less variation than real people.

A Nature study led by Ashwini Ashokkumar, covered by The Conversation, assembled 70 US experiments with close to 120,000 participants and asked GPT-4 to predict their outcomes. The predictions correlated strongly with the real results and ranked treatments well, but estimated effects at roughly twice their real size, and combining model forecasts with human forecasts beat either alone.

The threat runs the other way too. A paper in the Proceedings of the National Academy of Sciences built an autonomous synthetic respondent that passed 99.8 percent of 6,000 standard attention-check trials and could be instructed to skew a poll. A human panel that cannot tell its humans from a model has the same validity problem as a synthetic one. Verification cuts both ways.

Where Does a Synthetic Panel Need Real Respondents Behind It?

A synthetic panel needs real respondents in three places. They are the grounding data it is conditioned on, the validation sample its output is checked against, and the source of any answer the model has no basis to give. That last group includes new products, new markets and anything the training data did not see. Novelty is the gap.

That is the role Alchemic plays next to a synthetic study rather than against it. The interviews are AI-moderated, which is what makes the human sample fast enough to sit in the same workflow. A 200-interview qualitative study runs from brief to live dashboard in about three days, on WhatsApp with no link or app, by AI phone call, or on the web. The platform publishes 57+ languages including Hindi, Tamil and Telugu.

Respondents are recruited as managed fieldwork or from the brand's own list, with quality flags on every conversation.

Grounding is the strongest case. Each study feeds a knowledge base that persists across waves, so a twin built on a brand's own interviews, which is the design Bain recommends, has a first-party base that grows with every study rather than a third-party panel. The reliability guide for AI-moderated interviews covers when that base is sound enough to build on.

Validation is the case the codes require. A synthetic read of a new pack or a new market can be checked against 50 to 200 real conversations in the same week, and the sample validity guide covers how large that check needs to be. The human layer has fielded in the USA and the UK as well as across India, the Gulf and Southeast Asia, which matters for the subgroups where the arXiv study found synthetic error concentrates.

Where Synthetic Respondents Break

Synthetic respondents break on novelty, on subgroups, on magnitude and on anything that depends on lived experience. Each failure is documented. Each maps to a decision a brand should not make on synthetic data alone.

Novelty. A twin is conditioned on what people have already said. A product with no precedent, a claim the category has not made, or a market the training data underrepresents gives the model nothing to imitate. NIM's guidelines note that standard models are not designed to replicate lived experience or cultural nuance.

Subgroups. The arXiv study found distortions concentrated in underrepresented groups, and NIM's twins lost the variation between respondents. A synthetic read of a Tier 2 Indian shopper or a first-generation buyer is exactly where the error sits.

Magnitude. The Nature study's twofold overestimate and NIM's 1.2-point positivity bias mean a synthetic study can rank concepts and still misprice the winner. Purchase intent, price sensitivity and any go or no-go threshold need a human number. Ranking is not pricing.

Decisions with consequences. The Insights Association's guidance is that emerging methods be evaluated for validity before they inform decisions, and that reports distinguish human from synthetic findings. A launch, a pricing move or a claim substantiation built on a synthetic sample alone does not meet that bar. For those, a concept test with real shoppers, or personas built from real interviews, is the instrument the codes describe.

Frequently Asked Questions

What are synthetic respondents in market research?
Synthetic respondents are large language models answering survey questions or interview prompts as if they were people with a specified profile. Vendors sell three versions: a prompted persona with no real data behind it, a digital twin conditioned on real responses, and synthetic rows added to a human survey to enlarge a small segment. The 2025 ICC/ESOMAR Code defines synthetic data and requires its use to be disclosed.
Are synthetic respondents accurate?
Partly. Across 285 published silicon-to-human comparisons reviewed by the Nuremberg Institute for Market Decisions, 24.9 percent agreed, 65.3 percent diverged and 9.8 percent partially aligned. Twins do better on direction and rank than on magnitude: NIM's own experiment matched real choices 79 percent of the time while overrating brands by 1.2 points on a seven-point scale and compressing the spread of opinion.
What is the difference between a synthetic respondent and a digital twin?
A synthetic respondent is any model-generated answer standing in for a person. A digital twin is the grounded version: a model conditioned on real data from a specific population, a panel or an individual, so its answers track that source. Bain's May 2026 work found twins built on a company's own respondent-level data replicated about 90 percent of a prior conjoint study's outcomes; a prompted persona has no such base.
Can synthetic respondents replace a survey panel?
Not for decisions with consequences. Both the ICC/ESOMAR Code and the Insights Association's 2026 guidance require validation before synthetic findings inform decisions and disclosure of which findings are synthetic. Synthetic samples screen variants, pretest instruments and extend thin segments well. Launch, pricing, claims and any new market or product still need real respondents, either as the study or as the check.
How do you validate a synthetic panel?
Run the same instrument on a real sample and compare distributions, not just means. NIM's experiment used direct scale ratings and semantic similarity scoring against real respondents and reported match rates, bias by brand familiarity and variance. A validation sample of 50 to 200 real interviews in the same week catches the two documented failures, inflated positivity and compressed spread, before the synthetic read is used.

About the Author

Sreenadh Narayanan is the founder of Alchemic, an AI-powered consumer research platform used for ad testing, concept testing and brand tracking. He writes Alchemic's guides on qualitative research and research methods, covering interview design, sample sizes and how teams turn customer conversations into decisions.