Last updated: 19 August 2026
How much a discussion guide has to carry depends entirely on how much latitude the moderator has to leave it. A discussion guide is the ordered set of questions and probing instructions that governs an interview, and with a human it is a starting point. With a system that asks exactly what you wrote in the order you wrote it, the guide becomes the whole study.
With one that holds a genuinely dynamic conversation and keeps probing until an answer is usable, the guide stays a starting point, and the writing burden drops sharply.
Most guides are quietly worse than they read. A good moderator repairs them in flight: skipping a question the participant already answered, rephrasing one that landed badly, chasing something interesting that was never on the page. The Nielsen Norman Group's January 2026 test of AI interviewers, ten participants across two platforms, found both followed the script rather than the insight and declined to reframe weak questions.
Systems differ in how much latitude the moderator has to leave the guide, so check that before assuming the behavior.
Where the moderator cannot repair a guide in flight, that work has to happen before fieldwork, on the document. So the first question to settle is which kind of system you are writing for. Here is what the rigid case involves, which is the harder one to plan for.
What Does the Guide Have to Do Now?
On a rigid system, everything the moderator used to do in the room. That is the shift worth planning for, because it changes how long guide-writing takes. On a system that probes adaptively and can leave the guide, several of the rows below stop being your job and become the moderator's.
| Judgment | With a human moderator | With a rigid AI moderator |
|---|---|---|
| Skipping an already-answered question | Done instinctively | Must be written as a rule, or it gets asked twice |
| Rephrasing a confusing question | Done live | Never happens; write it right the first time |
| Deciding when to probe | Read from the room | Set per question, in advance |
| Following an unexpected answer | The best part of the job | Largely unavailable |
| Managing time across sections | Adjusted on the fly | Must be designed into length |
| Noticing discomfort | Read from face and tone | Not available at all |
Read the bottom two rows as constraints rather than gaps to engineer around. Some of this you cannot recover with better writing, which is a reason to keep certain studies human-moderated rather than a reason to write harder. Where bias actually enters such a study is mapped in AI moderator bias.
How Should You Write the Questions Themselves?
Write them the way you would say them out loud, to one person, with no chance to clarify. That single test catches most of what goes wrong.
- One idea per question. Double-barreled questions produce answers to whichever half the respondent noticed. A human moderator would catch that and split it live; write it as two questions so no moderator has to.
- Concrete over abstract. "Walk me through the last time you bought this" beats "what factors influence your purchase decisions". Recall of a specific event is more reliable than self-theorized behavior.
- Behavioral before attitudinal. Ask what happened before asking how they feel about it, so the account is not shaped by the opinion. Reversing this order is one of the most common and least noticed ways a guide manufactures its own findings.
- No embedded assumption. "What frustrated you about the checkout?" presumes frustration. "How did the checkout go?" does not.
- Plain vocabulary. Category jargon invites people to perform expertise rather than describe experience, and the performance is convincing enough to survive analysis.
None of this is new advice. AAPOR's best practices for survey research already asks that questions be specific and cover one concept at a time.
It also asks that they use words the target audience actually uses, and the qualitative methods literature has long held that interview questions should be open ended, neutral and free of leading language. The difference is that with an AI moderator these stop being good practice and become the only line of defense.
What Does a Good Question Look Like?
One worked example makes the gap concrete. "What do you like about the packaging?" assumes liking, offers no route to indifference, and will be asked to all two hundred participants exactly as written.
"Tell me what you noticed first when you picked it up" asks for an observation rather than a verdict, and leaves room for the answer to be nothing much. A human moderator would have softened the first version by interview three. A rigid moderator will not.
Where Do You Set the Probing Rules?
Per question, deliberately, and this is the part most teams skip. Tools differ here: several expose a probing level you set for each question, while others run a dynamic conversation and keep probing on their own until the answer is usable, which leaves you setting intent rather than rules.
Alchemic works the second way: from the client's brief, its researchers tailor the AI and the guide in the builder, so the buyer is not writing probing rules question by question.
That control is also a decision about where your findings will have depth. Probe hard on three questions and lightly on the rest, and the three become your themes regardless of what mattered to the participant.
A workable default:
- Deep probing on the two or three questions your decision actually rests on.
- One follow-up on anything where a one-word answer would be useless.
- No probing on warm-ups, screeners and factual questions, where extra follow-ups just add length.
- Explicit probe text where the topic is nuanced, rather than leaving the model to generate its own.
Distribute the depth to match the decision, not the hypothesis. Probing hardest on the thing you hope is true is how a guide becomes a leading instrument.
Write the probe text yourself wherever the topic has nuance. A generated follow-up will be grammatical and generic, and generic follow-ups produce generic answers. For concept and stimulus work in particular, the probe that matters is usually the one asking what the person expected before they saw it.
How Long Should the Interview Be?
Shorter than the equivalent human session. And designed to a length rather than trimmed to one. That same testing found interruptions, long pauses, repetitive questions and poor time management were common in AI-moderated sessions, so the interview will not self-correct if you overrun.
Practical shape: 8 to 12 substantive questions for a 20 to 25 minute session, with probing budget accounted for. Every question with deep probing effectively costs two or three.
Front-load what matters. If attention or connection degrades, the loss lands at the end, so nothing decision-critical should live there.
Section order carries a second cost that is easy to miss. Early questions prime later ones, so a guide that opens by asking about price will produce price-led answers to everything after it. With a human moderator that priming varies a little between sessions. With an AI moderator it is applied identically to the whole sample, which turns a mild ordering effect into a consistent one. Where a human moderator still wins is worked through in AI against human moderated interviews.
Who Will Actually Take the Interview?
The guide assumes a participant who can join, and mode decides who that is. This is a guide-writing question rather than a recruitment one, because vocabulary, length and stimulus all depend on it.
The ITU's Facts and Figures 2025 reports mobile broadband coverage as nearly universal while quality and affordability gaps persist, and counts 2.2 billion people still offline, most in low and middle income countries. A guide built for a 25-minute video session quietly excludes anyone without a stable connection and a private room.
Writing for other modes changes the document:
- Messaging-based interviews need shorter questions, no long preambles, and tolerance for voice notes as answers.
- Phone interviews cannot use visual stimulus, so anything shown has to be described in words that work read aloud.
- Asynchronous formats need each question to stand alone, since a participant may answer across a day.
Alchemic runs interviews natively inside WhatsApp with no link and no app to install, and AI phone interviews to any working number including feature phones. If a study will run there, write for it from the first draft. Adapting a video guide afterward rarely works. The same selection effect is examined in sample validity and who you miss.
How Do You Pressure-Test the Guide?
Test it before the sample is spent, using cheap passes that each catch a different failure.
- Read it aloud, end to end. Anything you stumble over, a respondent will too.
- Have someone answer it deliberately badly, with minimum-effort one-word replies. Then check whether the probing rules recover anything useful.
- Run three pilot interviews on your hardest segment, not your easiest, and read the raw transcripts rather than the synthesis.
- Mark every question that suggests its own answer. Adversarial reading catches leading phrasing that looks neutral to its author.
- Check the first sentence of each section for an assumption the participant has not yet confirmed.
- Have a native speaker read it in every language you will field, at the register your participants actually use, since segment vocabulary rarely survives literal translation.
The pilot transcripts are where the real signal is. Look for questions that produced identical answers across all three participants, which usually means the question told them what to say. Look also for questions everyone misunderstood in the same way, which points at vocabulary rather than logic.
Budget a full rewrite after the pilot rather than a tidy-up. Guides written for AI moderation tend to fail structurally rather than cosmetically, because the problem is usually an assumption baked into the order or the probing map rather than a badly worded sentence. Teams that plan for one revision cycle ship better studies than teams that treat the pilot as a formality.
What the Guide Cannot Fix
Being honest about the ceiling keeps the effort proportionate.
- The unanticipated finding stays out of reach. No amount of pre-written probing substitutes for a moderator noticing something nobody planned for. For genuine discovery these tools are not yet suited to semistructured, in-depth work, and that boundary is still moving.
- Nonverbal signals are thin. The interviewer presents no face, and outside video modes it reads none either.
- A perfect guide on the wrong sample is still wrong. Mode and recruitment decide who answers.
- Consistency amplifies whatever you wrote, including the flaws, identically across every participant.
- Consent and welfare are not guide problems but they are your problems. The ESOMAR code and guidelines require that disclosure be genuinely understood, and the Insights Association and AAPOR publish complementary standards.
Platform defaults make guide decisions on your behalf unless you override them, and two are worth setting deliberately. Guides for individual interviews work best at roughly six to eight primary questions, and question order changes the answers that follow it, which under a fixed script applies identically to everyone rather than varying between sessions.

