Last updated: 19 August 2026
The three are separated by one property: how much of the conversation is decided in advance. IVR decides everything, a voice bot decides most of it, and an AI interview decides the next question from what the respondent just said.
That sounds like a technical distinction and it is really a research one. It determines whether an open-ended answer is possible at all, and therefore whether the method can produce anything beyond the options you already thought of.
Vendors in all three categories now describe themselves with the same vocabulary, which is why buyers routinely shortlist a voice automation product against a research instrument and compare them on price.
What Is the Actual Difference Between Them?
Depth of response. IVR captures a keypress or a single word, a voice bot captures a constrained answer inside a scripted flow, and an AI interview captures an open answer and then probes it.
Everything else follows from that. IVR is cheap because it asks almost nothing. A voice bot handles a transaction because transactions have known shapes. An AI interview costs more per completed call because the output is qualitative rather than categorical.
The confusion is worth naming because it is expensive in both directions. Buying IVR when the study needed probing produces confident-looking data with no explanation attached. Buying an AI interview for a two-question satisfaction check is paying for depth nobody will read.
What Does IVR Do Well?
Very short, closed-ended data collection at volume and low cost. Press one for yes, two for no, rate this from one to five. It has been in production for decades and it works.
The strengths are real and specific. There is nothing to install, it reaches feature phones, it scales to large samples cheaply, and the output needs no coding because the respondent has already categorized themselves.
Best for: post-call satisfaction, single-metric tracking, yes-or-no eligibility screening, high-volume polling on settled questions.
Limitations: no probing, no clarification, and steep mid-call abandonment as soon as the menu gets deep. An IVR study that needs a why has already failed at design stage.
The screening use is documented. A study in Tropical Medicine and International Health evaluated interactive voice response as a way to increase the representativeness of rural respondents, reporting the yield and cost of using it to pre-screen numbers before a longer survey.
What Is a Voice Bot For?
Completing a transaction over the phone rather than learning something. Order status, appointment booking, payment reminders, first-line support triage. It handles natural speech but inside a flow whose destinations are known.
The distinction from an AI interview is intent rather than capability. A voice bot is measured on task completion. A research instrument is measured on whether the answer is true and complete, which is a different objective and produces different design choices.
Some voice-bot platforms now market survey functionality, and that is legitimate for structured questionnaires. It becomes a problem when the questionnaire has open-ended sections, because a system optimized for routing tends to treat an unexpected answer as a failure to classify rather than as the finding. The platform landscape for that channel is set out in AI phone interview platforms.
What Makes an AI Interview Different?
The follow-up. An AI interview poses further inquiries based on a previous answer, asks for further clarifications when an answer is too simplistic, or goes deeper into the inconsistency in the provided answer instead of just noting it.
That capability has documented limits. The Nielsen Norman Group's January 2026 test of AI interviewers, ten participants across two platforms, found the two platforms it tested followed the script rather than the insight, sticking to the guide and declining to reframe weak or irrelevant questions. Probing is real; open-ended discovery is not.
The evidence that the method works at scale is early but genuine. One published deployment ran an LLM-based telephone survey across a United States pilot of 75 participants and a Peru deployment of 2,739, administering open and closed questions and navigating branching logic without an interviewer roster. Gallup has begun publishing research on AI phone interviewing as a methodological question in its own right.
The category now has a published methodological record. A paper in JAMIA Open set out automated survey collection with LLM-based conversational agents, covering both the conversation and the extraction of structured answers from it. Where a human moderator still wins is worked through in AI against human moderated interviews.
Which One Fits Your Study?
Match the tool to the answer you need, not to the budget line. The table below sorts on what each method can actually produce rather than on how each is marketed.
| Method | Who decides the next question | Output | Best for | Where it fails |
|---|---|---|---|---|
| IVR | Fixed menu, decided entirely in advance | Keypress or single word | Single-metric tracking, eligibility screening, cheap high-volume polling | Anything needing a reason |
| Voice bot | Scripted flow with natural-speech input | Structured fields, task outcome | Transactions, support triage, short structured questionnaires | Unexpected answers get treated as classification failures |
| AI interview | Generated from the previous answer | Open text or audio, coded into themes | Structured qualitative work at volume, multilingual fieldwork | Genuine discovery, and anything needing visual stimulus |
| Human CATI | The interviewer, live | Open text, interviewer judgment included | Regulated work, sensitive topics, small specialist samples | Cost per complete at current response rates |
IVR is the right answer more often than research vendors admit. If the study is one metric tracked weekly across a large base, an AI interview adds cost and latency for depth nobody will read.
However, few programs stay single-metric forever. When the tracked number moves and the team needs the reason behind it, Alchemic's AI phone interviews carry that open-ended follow-up on the same working numbers, including feature phones, rather than forcing the study into a browser.
Who Can Each One Reach?
All four reach anyone with a working phone number, which is the point. Where they diverge is who stays on the line, and that is a function of how much the method asks of the respondent.
The ITU's Facts and Figures 2025 reports mobile broadband coverage as nearly universal while quality and affordability gaps persist, and counts 2.2 billion people still offline, most in low and middle income countries. Voice reaches that population when a browser session cannot.
What Changes With Language?
A menu can be recorded in ten languages cheaply because the script never varies. An interview cannot, because the follow-up has to be generated in the respondent's language and register, and in code-mixed speech that is a harder problem than translation.
This is where language claims stop being a count. Alchemic's AI phone research runs across 57+ languages including Spanish, Hindi, Tamil, Bangla, Arabic and Indonesian, with fieldwork managed across fourteen markets that include the USA and the UK. A named list is still not an evaluation, so ask for a raw transcript in the language you actually need. The same selection effect is examined in sample validity and who you miss.
How Do These Fit Into a Wider Study?
Rarely as the whole study. Voice methods usually carry one or two segments inside a design that reaches the rest another way, and treating any of them as the entire instrument is where budgets go wrong.
The common shape is a split by reachability. Phone or IVR carries the segments a link cannot reach, browser or messaging carries the rest, and the instrument stays constant so the cells remain comparable.
That comparability is the constraint people forget. Running the same questionnaire across three modes produces three slightly different response distributions, and the differences have to be acknowledged in the method note rather than averaged away silently.
Asynchronous messaging sits usefully between voice and browser, since an interview conducted inside WhatsApp needs no app install and no fixed slot. Where a study needs both breadth and depth, a managed program can field several modes against one guide instead of running each as a separate procurement.
What Should You Decide Before You Shortlist?
Whether the study needs a reason or a number. That single question eliminates two of the four methods immediately, and it is answerable before any vendor is contacted.
- If you need a number tracked over time, IVR or a structured voice flow is sufficient, and depth is a cost with no reader.
- If you need a reason, only an interview format produces one, whether automated or human.
- If the topic is sensitive, route that section to a person regardless of what the rest of the study uses.
- If visual stimulus is required, voice is out for that section entirely, and the decision is which other mode carries it.
Everything after those four is economics rather than method. How the voice portion sits beside the other interview modes in the design matters more than which voice vendor wins the comparison. The wider version of that interrogation is in 14 questions to ask a research vendor about reach.
Where All Three Fall Short
None of them is a general-purpose replacement for a human interviewer, and the honest limits are shared rather than distinctive.
- Response rates do not improve because the caller is automated. An unknown number remains an unknown number, whatever is on the other end.
- No visual stimulus in any of them. Packaging, creative and interface work cannot run by voice, whoever or whatever administers it.
- No nonverbal signal. A human interviewer hears hesitation and slows down. None of these three has an equivalent judgment.
- Distress goes unrecognized. Where a respondent may need a person, an automated system that cannot detect that should not be alone on the call.
- The evidence base is young. One large published deployment plus early institutional work is a starting point, not a settled literature.
Professional standards apply identically across all three. The ESOMAR code and guidelines govern consent and welfare regardless of who administers the call, and the Insights Association and AAPOR publish complementary guidance on disclosure and data quality.
What About Ongoing Programs?
Repeat measurement changes the calculus. A one-off study can absorb a mode that fits awkwardly, while a program running monthly compounds that mismatch across every wave.
For a tracker, the deciding property is consistency rather than depth. Whichever method is chosen has to behave the same way in wave eight as in wave one, which favors simpler instruments over richer ones when the two conflict.
That is why continuous brand tracking often pairs a stable quantitative spine with a smaller qualitative layer, rather than trying to make one voice method carry both jobs at once.
Accuracy is not evenly distributed across speakers either. Work in npj Digital Medicine found significantly higher transcription error rates for non-native English speakers when testing widely used speech recognition models, which is a sampling problem rather than a technical footnote.
How Do You Tell Them Apart in a Demo?
Give a deliberately awkward answer and watch what happens next. That single test separates the three faster than any feature list, because it exercises the only property that actually differs.
- Answer a question with something off-menu. IVR re-reads the options. A voice bot tries to route you. An AI interview asks what you meant.
- Give a one-word answer to an important question. Watch whether anything requests detail.
- Contradict yourself deliberately between two answers, and see whether the contradiction is pursued or recorded.
- Ask for the raw transcript, not the dashboard. Synthesis output demos well regardless of interview quality.
- Check where people abandoned, not just the completion rate. Drop-off location tells you about method friction.
Whether the wider study belongs on voice at all is a separate question from which voice method to buy, and it is usually better answered alongside the other interview modes available than in isolation.

