Last updated: 19 August 2026
Respondent reach is the set of people a research platform can actually get into a study, and it is what the questions below are designed to expose. The questions that predict whether an AI-moderated interview platform will work for you are about recruitment and reach, not about moderation or synthesis. Moderation has converged across the category and synthesis demos well regardless of quality. Neither determines whether the people you need can take part at all.
This is a familiar pattern in research procurement. Telephone research kept selling on call quality and sample size long after the thing that actually broke it was who still answered the phone.
Pew Research Center's own response rates fell to 6% by 2018, from around 9% a few years earlier. The demo covered the wrong variable. It is worth not repeating that.
| Group | Questions | What it protects against |
|---|---|---|
| Recruitment | 1 to 5 | Slow fieldwork, unusable samples, unfilled quotas |
| Interview mode and access | 6 to 10 | Silent exclusion of whole segments |
| Language | 11 to 12 | Fluent-looking transcripts that lost the meaning |
| Standards and data | 13 to 14 | Consent, disclosure and compliance exposure |
Below are 14 questions, grouped by what they protect you against. Each has a good answer and a warning sign. Most can be verified during a paid pilot rather than taken on trust.
Why Do Reach Questions Matter More Than Feature Questions?
Because features are increasingly uniform and reach is not. The Nielsen Norman Group tested AI interviewers with ten participants across two platforms in January 2026, and the limitations were shared across tools rather than unique to one. Some systems follow the script rather than the insight and will not reframe a weak question, and the experience can read as almost conversational but still unnatural. How much latitude the moderator has to leave the guide differs by tool, so ask.
Reach, by contrast, varies enormously between vendors and is rarely on the demo agenda. It also compounds. A moderation weakness costs you some depth in every interview. A reach failure costs you an entire segment, and you will not see it in the completion report.
Scale is the reason this matters more each year. DataReportal's Digital 2026 report puts more than 6 billion people online. The ITU still counts 2.2 billion offline, most of them in low and middle income countries. The same selection effect is examined in sample validity and who you miss.
What Should You Ask About Recruitment?
1. Do you recruit participants, or do I?
The single biggest practical difference between platforms. Some bring participants, some manage panels, some expect you to arrive with a list.
- Good answer: a clear statement of which, with named sources for the markets you care about.
- Warning sign: "both", with no detail on how sourcing works in your specific market.
2. Where do participants in my target market actually come from?
Panel provenance determines sample quality. A vendor reselling a general-purpose panel in a market it has never worked in is a different proposition from one running its own fieldwork there.
- Good answer: named panel partners, or an explanation of direct recruitment.
- Warning sign: "our global panel network", unelaborated.
3. How do you screen, and what is the screen-out rate?
Screening quality separates a usable sample from an expensive one. Ask for the ratio of screened to completed for a study like yours.
- Warning sign: no screen-out figures available. It usually means nobody measures them.
4. How are incentives delivered in my market?
This is a real operational constraint and a common failure point. Bank transfers, mobile wallets, gift codes and local rails are not equally available everywhere, and incentive friction shows up as skewed completion. The consent and disclosure mechanics are covered in consent and disclosure in AI-moderated research.
5. What happens if you cannot fill a quota?
The honest answer involves a conversation and a timeline, not silent substitution. Where that sample comes from in the first place is covered in survey panels and where respondents come from.
- Warning sign: any suggestion that quotas are loosened without telling you.
What Should You Ask About Interview Mode?
6. What interview modes do you support?
Mode determines the sampling frame more than any other variable. Browser video, browser voice, messaging apps and telephone each reach a materially different population.
7. Can you interview someone without a stable broadband connection?
Pew Research Center's mobile technology fact sheet reports 16 percent of US adults as smartphone-only internet users, rising to 34 percent in households under $30,000 a year, and the ITU's Facts and Figures 2025 shows a steeper gradient across markets. If your study is browser-video-only, that gap becomes your sample boundary.
- Good answer: a specific non-browser mode, described concretely.
- Warning sign: "our platform is very lightweight."
8. Can a participant take part without installing anything or clicking a link?
Install friction and link friction both suppress participation, and they suppress it unevenly across segments.
9. Is participation synchronous or asynchronous?
Fixed live slots exclude shift workers, caregivers, small traders and anyone with intermittent connectivity. Asynchronous participation widens the frame considerably.
10. What device data do you capture, and can I see it?
You need this to audit your own coverage after the fact. A vendor that cannot show you the device and connection profile of your completions cannot help you check who you missed.
What Should You Ask About Language and Standards?
11. How many languages do you moderate in, and is moderation tuned or translated?
The advertised number matters far less than the distinction. Multilingual interviewing without translators is a genuine strength of AI moderation, but it only holds if the moderation actually works in the language.
- Good answer: a clear description of how language performance was evaluated.
- Warning sign: a large language count with no explanation of how any of them were tested.
12. Can I see a raw transcript in my target language?
The most useful single request in any vendor evaluation. Ask for an unedited transcript, not the platform's English translation of one, and have a native speaker read it. Fluency problems, register mistakes and missed idiom are immediately obvious and almost never appear in a demo.
13. Which research standards do you operate under?
The ESOMAR code and guidelines govern consent, participant welfare and data handling across the profession, and the Insights Association and AAPOR publish complementary standards. A vendor should be able to answer this without checking.
14. How is participant consent handled for AI moderation specifically?
Participants should know they are speaking with an AI system, and how their recording will be used. Disclosure wording is itself a design decision that affects how participants answer, so ask to see the exact text and where it appears in the flow, not a description of it.
How Should You Structure the Pilot?
Run one paid pilot against your hardest audience rather than your easiest, because an easy audience tells you nothing you did not already assume. The goal is not to confirm the platform works. It is to find the boundary of where it stops working, while you still have the option not to sign.
Judge what comes back against ordinary standards for conducting qualitative research rather than against the vendor's own dashboard, which is built to look convincing.
- Pick the segment you most often struggle to recruit. If the pilot reaches them, the rest is straightforward.
- Use a real guide from a real project, not a simplified test script. Weak questions are part of what you are testing.
- Set an explicit quota and measure completion against it by segment, not in total.
- Request the raw transcripts at the end, including from incomplete sessions. Abandoned interviews are the most informative artifact a pilot produces.
- Agree in advance what a failed pilot looks like. Teams that skip this reliably rationalize a disappointing result into a qualified success.
Where a vendor's model is software you operate rather than fieldwork run on your behalf, add one more check: whoever will actually run studies day to day should run the pilot, not the research lead who is evaluating it. Tools that are pleasant to evaluate and painful to operate are common in this category. The wider buying decision is covered in platforms that run the study for you.
Which Answers Are Hardest to Fake?
Three of them. Whether a participant can take part without a stable broadband session, whether participation can be asynchronous, and whether you can read a raw non-English transcript. Those three are difficult for a browser-only platform to answer well, which makes them the most efficient use of a first call.
| Capability | What it requires to be real | Why it is hard to retrofit |
|---|---|---|
| Non-browser interviews | Messaging or telephony infrastructure | Not a feature flag, it is a different delivery stack |
| Asynchronous participation | Session state held over hours or days | Conflicts with live-session product design |
| Tuned non-English moderation | Evaluation data per language | Requires sustained work in that market |
| Managed fieldwork | People on the ground, incentive rails | A services capability, not a software one |
Weight these four heavily, because they cannot be added between your first call and your contract. Feature gaps close in a quarter. Delivery-model gaps do not. The Alchemic and Listen Labs comparison sets out how recruitment models differ in practice.
Alchemic answers these by running interviews natively inside WhatsApp, with no link and no install. It adds AI phone interviews to any working number and moderation across 57+ languages including Spanish, Hindi, Tamil, Bangla, Arabic and Indonesian. Fieldwork is managed across fourteen markets, from the USA and the UK to South and Southeast Asia, the Gulf and Africa, rather than handed over as software.
Whether that fits depends entirely on where your respondents are. For a US-only study of high-connectivity consumers, a browser platform may answer every question above perfectly well.
What These Questions Will Not Tell You
Reach due diligence has clear limits, and it is worth being explicit about them so the checklist is used well rather than used alone.
- It says nothing about interview quality. A platform can have excellent reach and a mediocre moderator. Read transcripts as well as asking these questions.
- It will not fix a weak discussion guide. A rigid AI interviewer does not reframe weak or irrelevant questions, so the upstream fix is question design, where the methods literature on open-ended interview questions is worth more than any vendor feature. Reach multiplies whatever your guide already is.
- Procurement timing matters. Reach questions are cheapest to ask before a shortlist forms, because once a preferred vendor exists these questions get reframed as objections to overcome rather than as criteria to evaluate against.
- Vendor answers are claims until tested. Most of these should be verified in a paid pilot with your actual audience, not accepted on a call.
- Not every study needs maximum reach. For a premium product in a high-connectivity market, the browser frame may be close to your real market. The point is to choose it deliberately.
- Some work still needs human moderation. For high-stakes decisions, sensitive topics and open discovery, independent testing found human interviewers still outperform, whatever the reach.
The best evaluations combine this list with a small paid pilot and a transcript review. Industry trade coverage is useful background on how vendors are positioning themselves as the category consolidates, but positioning is not evaluation and should not be read as one.

