Last updated: 19 August 2026
CATI, CAPI and CAWI are three ways of administering the same questionnaire: by phone, in person or online. The questionnaire can be identical in all three and still produce different answers, which is why the choice is a research decision rather than a procurement one.
Most debates about data collection modes revolve around cost per complete, which is the least interesting variable, because it barely separates well-designed studies from badly designed ones.
What actually differs is who you can reach, what you can show them, and how honestly they answer.
What Do CATI, CAPI and CAWI Actually Mean?
Three computer-assisted ways of running the same interview: by telephone, in person, and on the web. The software carries the questionnaire logic in all three; the mode is how you reach the respondent.
- CATI is computer-assisted telephone interviewing. An interviewer calls, reads questions from a screen, and enters answers as the respondent speaks.
- CAPI is computer-assisted personal interviewing. An interviewer is physically present, usually with a tablet, and can show stimulus and observe the respondent directly.
- CAWI is computer-assisted web interviewing. The respondent completes a questionnaire themselves in a browser, with no interviewer present at all.
The shared "computer-assisted" part means routing, piping and validation are handled by software. That has been true of all three modes for decades, so it is not what separates them. What replaced the traditional phone room is covered in AI CATI software replacing phone interviews.
What matters is whether an interviewer is present and whether the respondent can see anything. The platform landscape for that channel is set out in AI phone interview platforms.
Which Mode Should You Use?
Start from what the study needs to find out, not from the budget. Three questions settle most cases: does the respondent need to see anything, is the topic one people shade the truth on, and can the channel reach the population at all?
| Factor | CATI (phone) | CAPI (in person) | CAWI (web) |
|---|---|---|---|
| Interviewer present | Yes, by voice | Yes, in person | No |
| Can show stimulus | No | Yes, fully | Yes, on screen |
| Typical reach | Anyone with a working phone | Geographically bounded | Anyone online and willing |
| Cost per complete | Middle | Highest | Lowest |
| Speed to field | Fast | Slowest | Fastest |
| Sensitive topics | Weaker | Weakest | Strongest |
| Complex questions | Weaker | Strongest | Middle |
| Open-ended depth | Good | Best | Weakest |
| Main failure mode | Non-response | Cost and coverage | Self-selection and inattention |
The table is a starting point rather than a verdict. A study with on-screen stimulus and a sensitive topic has a genuine conflict between rows, and resolving that conflict is the actual design work.
Why Does the Same Question Get Different Answers by Mode?
Because a person answering a stranger's voice behaves differently from a person answering a form. This is called a mode effect, and it is measurable rather than theoretical.
Two mechanisms drive most of it. Social desirability pushes answers toward what the respondent thinks the interviewer wants to hear, and it is strongest when a human is listening. Research summarized in the medical and social science literature on sensitive-question methods documents this effect consistently across topics.
The second is that visual and aural presentation change how scales are used. A respondent hearing five options remembers the last ones best; a respondent seeing them treats the list spatially.
Pew Research Center's methodological work found mode-of-interview effects when moving from telephone to the web, with differences concentrated in questions about personal wellbeing and social attitudes rather than uniformly across the questionnaire. Separately, Pew found few mode effects on questions about news consumption habits, which is the useful counterweight: mode effects are real but topic-dependent.
What Does That Mean for Tracking Studies?
That you cannot switch modes mid-series and read the resulting movement as market change. A mode switch introduces a discontinuity that looks exactly like a trend.
The standard handling is a parallel run: field both modes simultaneously for one wave, measure the gap on each key metric, and carry the gap forward as a documented adjustment. Skipping the parallel run does not avoid the problem, it only hides it. Where continuity matters, brand tracking designed around the mode change is cheaper than re-baselining a whole series later.
A 2026 study in the Journal of Medical Internet Research measured this on a live instrument, examining mode effects between mobile web and telephone surveys on patient experience scores and finding differences present rather than negligible. Where bias actually enters such a study is mapped in AI moderator bias.
When Is CAPI Worth Its Cost?
When the study needs physical stimulus, observed behavior, or a population that cannot be reached any other way. Those three cases justify the expense; familiarity does not.
Physical stimulus is the easiest case. A pack test cannot run over the phone, and a photo is not equivalent to holding the object.
The second is observation. Watching what a respondent actually does, rather than what they report doing, is a difference no other mode captures.
The third is coverage. In segments or markets where neither internet nor phone reliably reaches the population, in-person fieldwork is not the premium option but the only one that produces a sample.
Where Does CAWI Break Down?
On engagement and on who volunteers. A web questionnaire has no interviewer to notice that the respondent stopped reading on question nine, and no way to tell a considered answer from a fast one.
- Straightlining and speeding are invisible without deliberate quality checks built into the instrument.
- Open-ended answers collapse. Typing effort suppresses length, so the qualitative portion of a web survey is usually its weakest section.
- Self-selection skews the sample toward people who enjoy taking surveys, which is not a neutral trait.
- Panel fatigue compounds it, since the same professional respondents recur across studies.
- Device truncation cuts long grids on mobile, where most completes now happen.
None of that makes CAWI a poor method. It means CAWI's main risk is engagement quality where CATI's is reach, and the two failure modes need different defenses.
What Is Replacing the Traditional Three-Mode Choice?
Voice-based automation, which complements all three modes rather than substituting for any of them. It keeps the reach of a phone line while removing the per-hour interviewer cost that made CATI expensive at depth.
The evidence base is early but real. A study published on arXiv evaluated an LLM-based telephone survey system at scale, running it against human interviewers and an online form. Gallup has announced its own research program into AI phone interviewing, which matters mainly because Gallup's methodological caution is well established.
AI phone interviews inherit CATI's coverage property, since they reach any working handset including feature phones, while removing the per-hour interviewer cost that historically capped sample sizes. What they do not inherit is a human interviewer's judgment, which is a genuine trade rather than a free upgrade.
The Nielsen Norman Group's January 2026 test of AI interviewers, ten participants across two platforms, found the two platforms it tested followed the guide rather than the insight, sticking to the script and declining to reframe weak questions. That is a limitation for exploratory qualitative work and much less of one for a structured instrument.
How Do Open-Ended Questions Perform Across Modes?
Poorly online, fairly well by telephone, and best face to face. That is the largest quality difference between the three modes, and the one most often overlooked when cost drives the decision.
The reason is effort. Typing a considered paragraph into a browser form is work, so respondents write less than they would say.
Telephone removes the effort barrier and introduces a different one. The interviewer summarizes while the respondent speaks, so what gets recorded is not always the respondent's own wording.
Face to face, an interviewer can probe an unclear answer in the moment, and that follow-up is where the most valuable material surfaces. It is also the most expensive minute in market research.
Can Automation Close the Open-Ended Gap?
Partly, and only where the follow-up is genuinely adaptive. An automated interviewer that reads the answer, notices what was left unsaid, and asks about it recovers some of what a web form loses.
The limit is not mechanical. AI-moderated interviews can probe from several angles for every respondent, where a human interviewer six hours into a shift cannot, but no automated probe fixes a faulty question.
The result is that automation raises the floor of quality rather than the ceiling. A bad guide produces consistently bad output, which is easier to detect and harder to excuse than uneven human fieldwork.
How Do You Combine Modes Without Corrupting the Data?
Assign modes to populations deliberately and record the assignment. A mixed-mode study is defensible when the mixing is purposeful and documented, and indefensible when it simply reflects whoever was available to field it.
- Split by reachability, not by convenience, so each segment gets the mode that actually reaches it.
- Keep question wording identical across modes, even where one mode tempts you to shorten.
- Record mode as a variable on every response so mode effects can be tested rather than assumed.
- Run a parallel wave whenever a track changes mode, and publish the measured gap.
- Report the mode split in the deliverable, since a reader cannot interpret the numbers without it.
Professional standards support this directly. The ESOMAR code and guidelines and AAPOR's standards and ethics both treat method disclosure as a baseline obligation rather than an optional appendix.
For studies that span several countries, mode availability itself varies by market, and scoping fieldwork before entering a new market usually settles the mode question faster than a global template does. Where a program spans qualitative and quantitative work at once, managed end-to-end delivery keeps the mode decision with the people who will have to defend the numbers. When that data holds up is examined in when AI-moderated interviews produce reliable data.
Honest Limits of Any Mode Comparison
A table cannot decide a study, and treating it as though it can produces confident errors.
- Cost per complete is not cost per usable answer. The cheapest mode is expensive if half the opens are inattentive.
- Reach figures are national averages. ITU's Facts and Figures 2025 reports near-universal mobile broadband coverage alongside 2.2 billion people still offline, and your segment may sit on either side of that.
- Mode effects are topic-specific. Assuming a uniform correction across a questionnaire is its own error.
- Panel quality varies more than mode does. A good CAWI panel beats a bad CATI sample on almost every dimension.
- Speed claims assume recruitment is solved. Fielding is fast; finding the right people rarely is.
These trade-offs are not fixed, and they are not moving at the same speed. The voice rows are changing fastest.

