Last updated: 18 September 2026
Quick Answer: A qualitative questionnaire is a self-completion instrument of open-ended questions respondents answer alone, with no moderator to follow up. Build one from a screening intent, a warm-up, core open items, written probe substitutes, and a close, with the item you most need answered second, not last.
A 2018 PLOS ONE analysis of 28 interview datasets covering 1,147 interviews found ten-person samples surfaced 95 percent of a domain's salient ideas exhaustively, against 53 percent capped at three, matching a single-pass form.
What Makes a Qualitative Questionnaire Effective Without a Moderator?
An effective qualitative questionnaire answers one research question, pretested on people like the sample. No moderator is present to catch a misreading: in a 2017 web-probing test of one interview item, just 5 percent of online respondents flagged an unfamiliar term, against 30 percent face to face.
Five devices substitute for the missing follow-up, at least two per instrument:
- The piped follow-up quotes the respondent's own words, then asks what was behind them. Because the prompt is the respondent's own sentence rather than the researcher's paraphrase, there is nothing to dispute before they answer the real question.
- The paired opposite asks for both valences in one item. Asked as a single "what did you think," the item collects whichever valence is top of mind; asking for both in the same breath stops one from crowding out the other.
- The specificity demand asks for a moment, never a general view. A single remembered instance is harder to answer with a rehearsed opinion than a general attitude question, and it gives the coder something concrete to place in a code rather than a restated sentiment.
- The falsification item asks what would change their mind. It is the one item built to surface an objection a satisfied-sounding respondent has no other reason to volunteer.
- The catch-all close recovers whatever else was missed. It carries no fixed target, so it is usually where a theme the codebook has not yet named shows up first.
None of the five recreates a live follow-up exactly; used together, at least two per instrument, they cover most of what a moderator sitting in the room would have asked next.
Pretesting is mandatory: Statistical Quality Standard A2 requires it, and AAPOR's best practices call for cognitive interviews to catch misreadings.
Qualitative Questionnaire Examples by Study Type
Three instruments follow, one each for customer experience, concept reaction and employee experience, consent secured before item 1 per the ICC/ESOMAR Code. The most valuable item also goes early in each: 68 percent answered it near the start against 24 percent near the end on the same long-form survey.
A Customer Experience Questionnaire
Screening intent: bought in the last 30 days, used it twice, not brand- or rival-employed.
- What made you start looking, the day you decided to buy?
- Walk us through the first time you used it, in order.
- Name one thing that worked better than expected, one that worked worse.
- You mentioned something that worked worse. Describe that moment.
- If a friend asked whether to buy it, what would you say?
- How did the price feel once you saw it, against what you expected to pay?
- Did you consider anything else before choosing this, and what tipped the decision?
- Anything about this purchase we did not ask that matters?
Item 3 is the paired opposite, item 4 the piped follow-up, item 2 the sequence. Item 7 pulls in the alternatives actually considered, which a straight satisfaction score never surfaces.
A Concept Reaction Questionnaire
Screening intent: category buyer in the last three months, new to this concept.
- In your own words, what is this product and who is it for?
- What, if anything, was unclear about what you just read?
- What was the first thing you thought? Write it exactly as it came.
- On a scale of 1 to 5, how likely would you be to buy this?
- What would make you doubt the main claim?
- What would you expect this to cost, before you are told the price?
- What, if anything, would this replace for you?
- Anything about the concept we have not asked that matters?
Items 1 and 2 check comprehension before item 3's reaction; item 5 recovers the objection a polite 4 conceals. Item 6 catches a price anchor before the real number reframes it; item 7 places the concept against whatever it would actually displace.
An Employee Experience Questionnaire
Internal research carries a risk the other two instruments do not: the respondent and the person who reads the results often share a reporting line, so wording has to protect against blame rather than invite it. A question that would read as neutral in a customer study can read as a performance review in an employee study, and respondents answer accordingly: guarded, brief, and safely positive. The items below route around that by asking for a process or a specific moment rather than a rating of a person.
Screening intent: current employee, at least six months' tenure, not the people-manager of the team being studied, since a manager in the sampling frame changes what the results can honestly claim to represent.
- Think back to a recent week that felt genuinely good at work. What made it good?
- Walk through what your first two weeks in this role actually looked like.
- Name one part of your role that is easier than it should be, and one that is harder.
- You mentioned something that is harder than it should be. What gets in the way, specifically?
- If a friend asked whether to take a job here, what would you tell them?
- What would have to change for you to start looking for another job?
- Which tool, process or approval step costs you the most time for the least value?
- Anything about working here we did not ask that matters?
Item 3 is the paired opposite, item 4 the piped follow-up, item 6 the falsification item aimed at retention instead of a purchase. Item 7 asks about a process, not a person, which keeps a friction complaint out of performance-review territory. Anonymity has to be a written guarantee here, not an assumption: state who reads raw responses before item 1, not after someone has already answered item 4.
How Do You Write Questions That Produce Depth Instead of Length?
Depth comes from narrowing the question: one episode, a timeframe, the sequence.
Pew Research Center found nonresponse on open-ended questions runs near 18 percent against 1 to 2 percent for closed ones, and places them near the beginning since an earlier closed item primes the answer.
- Episode items, not opinion items. Describe the last onboarding mishap, not an opinion of onboarding.
- One question per item. Mixing ease and value in one item muddles the answer.
- Scale after story. A rating primes its own justification, and question order changes what a group will say in moderated work too; a form cannot rewind.
Length compounds the same problem. Pew's own conclusion does not blame one question in isolation; it names the number of questions in the instrument, open and closed together, as one of the forces behind any single item's refusal rate. The placement data above cuts the same way inside one long form: a near-beginning open item scored 68 percent, the near-end item on the identical survey scored 24 percent, even though both were optional and both reached a respondent who had already finished the rest of the questionnaire. Every open item added to an instrument is a tax the later items pay, not just the one in front of the respondent right now, which is why the FAQ below puts a number on where that tax gets expensive.
Questionnaire, Interview or Both: Which Fits the Question?
A questionnaire fits a settled question; an interview fits an unknown probe. Both fit in sequence, questionnaire first, interview after to explain the surprise, since a questionnaire is cheaper to run at volume but has no way to ask a respondent what an unexpected answer meant.
Rows are sorted alphabetically by instrument.
| Instrument | Follow-Up | Reaches | Best For |
|---|---|---|---|
| AI phone interview | Adaptive, spoken, in the moment | Any phone, no smartphone | Low typing comfort, weak data |
| Live moderated interview | Adaptive; can redirect mid-session | Can schedule, attend | Exploratory work, unknown probe |
| Self-completion questionnaire | None; every probe written in advance | Literate, typing-comfortable, in-language | Breadth, sensitive topics, settled questions |
| Voice-note questionnaire | Written probes; answers arrive as speech | Talk, not type | Long answers a text box would silence |
| Auto-follow-up questionnaire | Conditional, on length | Sits longer | Near-survey depth |
Research platforms combining qualitative and quantitative studies earn their price when the number and reason share a respondent; Alchemic's AI-moderated interviews carry open probing and structured items in one session, so a quote bank and distribution chart share a field.
Quantitative questionnaire examples serve a different job, comparability over meaning; the qualitative research pillar covers it.
How Do You Analyze Qualitative Questionnaire Responses?
Code text against a codebook refined early, before fielding: saturation runs a median 75 interviews when exhaustive, 16 when capped at three. A 2026 arXiv study had five coders and an LLM code 903 open-ended answers across six variables, reaching real alignment: an Adjusted Rand Index of 0.61 for coding, 0.54 for theme generation.
- Codebook control. You define the codes; no fixed taxonomy is imposed.
- Traceability. Every theme traces back to its source response.
- Coding in the source language. Translating first loses the distinctions worth studying.
- Stable codes across waves. Codes that shift between waves cannot show change.
Building and Applying the Codebook
A codebook turns a code like "frustration" from a private judgment call into something two coders would apply the same way. A BMC Medical Research Methodology case study of codebook development gives each code a label, a working definition, a description of what counts, explicit exclusions, and a verbatim example pulled straight from a response. A single entry built that way might read: label, "response friction"; definition, a moment the respondent had to work harder than expected to get value; description, covers delays, unclear instructions and repeated contact for the same issue; exclusion, not a price complaint, which gets its own code; example, the verbatim pulled from item 4 of the Customer Experience instrument above. Written out this way, two coders reading the same response reach for the same code instead of two different words for the same idea.
The build then runs in passes, not one read-through:
- A first pass on a sample, not the full set. Before the codebook is trusted against everything, it runs against a subset: responses get margin notes on candidate codable moments, checked against a draft built from the research question and an initial scan of the data.
- The draft codebook applies to the wider set. Coding continues in rounds; a code that does not fit gets refined, and a moment that fits nothing gets a new code. The same case study treated the codebook as stable once a further round produced nothing new.
- Reliability gets tested, not assumed. Two kinds of agreement get checked: whether one coder reaches the same call on a second look, and whether two coders reach the same call on the same response, scored as agreements divided by agreements plus disagreements. A rule-of-thumb floor of 75 percent counts as adequate; below it, the disagreement gets discussed and the code definition gets sharpened, not the response thrown out.
- Frequency gets reported next to the verbatim, never alone. Take item 4 of the Customer Experience instrument again: a first-pass sample might sort "worked worse" answers into a handful of repeatable codes, a shipping-delay code, a packaging code, a support-response code. The write-up pairs each code's count with the verbatim that captures it best: nine of forty responses might code to shipping delay, reported next to the one line that explains why. The count shows the pattern is real; the quote is what a reader remembers.
None of this requires more than one coder, but it is what makes a second coder's check meaningful once a finding has to carry weight beyond the team that ran the study.
The same holds for forms, chats, transcripts, per open-end coding software.
Which Respondents a Written Questionnaire Never Reaches
A written questionnaire filters its sample three ways: literacy, typing comfort, language, before any screener runs.
The 2023 adult skills assessment found 28 percent of US adults aged 16 to 65 at the lowest literacy level or below, up from 19 percent in 2017.
Voice notes mean speech, not typing, and count as usable qualitative data in-language; phone delivery reaches people with no smartphone. Alchemic runs both: WhatsApp-native interviews need no link or app. It publishes 57+ languages including Spanish, Arabic and Mandarin, with managed fieldwork or bring your own panel across 14 markets including the USA and the UK.
Where a Qualitative Questionnaire Breaks Down
The form fails four ways; three produce data that looks fine:
- Selective nonresponse. Respondents who answer open items skew lower income and satisfaction than those who skip them.
- No repair path. A misread item costs thirty seconds live; in a questionnaire it costs every respondent.
- Automated probing is not settled. A 2025 review weighs whether AI-driven probing can substitute for a live prober; chat formats demand more effort, with limited evidence of sustained engagement, and latitude varies by tool.
- The unasked question. A catch-all close recovers what an instrument missed. It does not recover a question nobody framed, which is the job of exploratory interviews.
Once it has the client's brief, Alchemic tailors the discussion guide to handle these risks before fielding, rather than leaving that work to the buyer. An unpretested questionnaire is unmeasured, not cheaper.

