Home Feeds Careers Get in Touch

How to Write a Discussion Guide for an AI Moderator

Aug 25, 2026Sreenadh NarayananSreenadh Narayanan9 min read
discussion guide ai moderator interview guide writing ai moderated interview questions research guide design probing rules ai interview qualitative interview guide how to write interview questions moderator guide template
Writing a discussion guide for an AI-moderated interview

TL;DR

  • A human moderator repairs a weak guide live, so most guides read better than they perform
  • How much repair falls on you depends on the system: a rigid moderator asks exactly what you wrote, a dynamic one keeps probing until answers are usable
  • Writing for the rigid case means moving every live judgment into the document first

Last updated: 19 August 2026

How much a discussion guide has to carry depends entirely on how much latitude the moderator has to leave it. A discussion guide is the ordered set of questions and probing instructions that governs an interview, and with a human it is a starting point. With a system that asks exactly what you wrote in the order you wrote it, the guide becomes the whole study.

With one that holds a genuinely dynamic conversation and keeps probing until an answer is usable, the guide stays a starting point, and the writing burden drops sharply.

Most guides are quietly worse than they read. A good moderator repairs them in flight: skipping a question the participant already answered, rephrasing one that landed badly, chasing something interesting that was never on the page. The Nielsen Norman Group's January 2026 test of AI interviewers, ten participants across two platforms, found both followed the script rather than the insight and declined to reframe weak questions.

Systems differ in how much latitude the moderator has to leave the guide, so check that before assuming the behavior.

Where the moderator cannot repair a guide in flight, that work has to happen before fieldwork, on the document. So the first question to settle is which kind of system you are writing for. Here is what the rigid case involves, which is the harder one to plan for.

What Does the Guide Have to Do Now?

On a rigid system, everything the moderator used to do in the room. That is the shift worth planning for, because it changes how long guide-writing takes. On a system that probes adaptively and can leave the guide, several of the rows below stop being your job and become the moderator's.

Judgment With a human moderator With a rigid AI moderator
Skipping an already-answered question Done instinctively Must be written as a rule, or it gets asked twice
Rephrasing a confusing question Done live Never happens; write it right the first time
Deciding when to probe Read from the room Set per question, in advance
Following an unexpected answer The best part of the job Largely unavailable
Managing time across sections Adjusted on the fly Must be designed into length
Noticing discomfort Read from face and tone Not available at all

Read the bottom two rows as constraints rather than gaps to engineer around. Some of this you cannot recover with better writing, which is a reason to keep certain studies human-moderated rather than a reason to write harder. Where bias actually enters such a study is mapped in AI moderator bias.

How Should You Write the Questions Themselves?

Write them the way you would say them out loud, to one person, with no chance to clarify. That single test catches most of what goes wrong.

  • One idea per question. Double-barreled questions produce answers to whichever half the respondent noticed. A human moderator would catch that and split it live; write it as two questions so no moderator has to.
  • Concrete over abstract. "Walk me through the last time you bought this" beats "what factors influence your purchase decisions". Recall of a specific event is more reliable than self-theorized behavior.
  • Behavioral before attitudinal. Ask what happened before asking how they feel about it, so the account is not shaped by the opinion. Reversing this order is one of the most common and least noticed ways a guide manufactures its own findings.
  • No embedded assumption. "What frustrated you about the checkout?" presumes frustration. "How did the checkout go?" does not.
  • Plain vocabulary. Category jargon invites people to perform expertise rather than describe experience, and the performance is convincing enough to survive analysis.

None of this is new advice. AAPOR's best practices for survey research already asks that questions be specific and cover one concept at a time.

It also asks that they use words the target audience actually uses, and the qualitative methods literature has long held that interview questions should be open ended, neutral and free of leading language. The difference is that with an AI moderator these stop being good practice and become the only line of defense.

What Does a Good Question Look Like?

One worked example makes the gap concrete. "What do you like about the packaging?" assumes liking, offers no route to indifference, and will be asked to all two hundred participants exactly as written.

"Tell me what you noticed first when you picked it up" asks for an observation rather than a verdict, and leaves room for the answer to be nothing much. A human moderator would have softened the first version by interview three. A rigid moderator will not.

Where Do You Set the Probing Rules?

Per question, deliberately, and this is the part most teams skip. Tools differ here: several expose a probing level you set for each question, while others run a dynamic conversation and keep probing on their own until the answer is usable, which leaves you setting intent rather than rules.

Alchemic works the second way: from the client's brief, its researchers tailor the AI and the guide in the builder, so the buyer is not writing probing rules question by question.

That control is also a decision about where your findings will have depth. Probe hard on three questions and lightly on the rest, and the three become your themes regardless of what mattered to the participant.

A workable default:

  • Deep probing on the two or three questions your decision actually rests on.
  • One follow-up on anything where a one-word answer would be useless.
  • No probing on warm-ups, screeners and factual questions, where extra follow-ups just add length.
  • Explicit probe text where the topic is nuanced, rather than leaving the model to generate its own.

Distribute the depth to match the decision, not the hypothesis. Probing hardest on the thing you hope is true is how a guide becomes a leading instrument.

Write the probe text yourself wherever the topic has nuance. A generated follow-up will be grammatical and generic, and generic follow-ups produce generic answers. For concept and stimulus work in particular, the probe that matters is usually the one asking what the person expected before they saw it.

How Long Should the Interview Be?

Shorter than the equivalent human session. And designed to a length rather than trimmed to one. That same testing found interruptions, long pauses, repetitive questions and poor time management were common in AI-moderated sessions, so the interview will not self-correct if you overrun.

Practical shape: 8 to 12 substantive questions for a 20 to 25 minute session, with probing budget accounted for. Every question with deep probing effectively costs two or three.

Front-load what matters. If attention or connection degrades, the loss lands at the end, so nothing decision-critical should live there.

Section order carries a second cost that is easy to miss. Early questions prime later ones, so a guide that opens by asking about price will produce price-led answers to everything after it. With a human moderator that priming varies a little between sessions. With an AI moderator it is applied identically to the whole sample, which turns a mild ordering effect into a consistent one. Where a human moderator still wins is worked through in AI against human moderated interviews.

Who Will Actually Take the Interview?

The guide assumes a participant who can join, and mode decides who that is. This is a guide-writing question rather than a recruitment one, because vocabulary, length and stimulus all depend on it.

The ITU's Facts and Figures 2025 reports mobile broadband coverage as nearly universal while quality and affordability gaps persist, and counts 2.2 billion people still offline, most in low and middle income countries. A guide built for a 25-minute video session quietly excludes anyone without a stable connection and a private room.

Writing for other modes changes the document:

  • Messaging-based interviews need shorter questions, no long preambles, and tolerance for voice notes as answers.
  • Phone interviews cannot use visual stimulus, so anything shown has to be described in words that work read aloud.
  • Asynchronous formats need each question to stand alone, since a participant may answer across a day.

Alchemic runs interviews natively inside WhatsApp with no link and no app to install, and AI phone interviews to any working number including feature phones. If a study will run there, write for it from the first draft. Adapting a video guide afterward rarely works. The same selection effect is examined in sample validity and who you miss.

How Do You Pressure-Test the Guide?

Test it before the sample is spent, using cheap passes that each catch a different failure.

  • Read it aloud, end to end. Anything you stumble over, a respondent will too.
  • Have someone answer it deliberately badly, with minimum-effort one-word replies. Then check whether the probing rules recover anything useful.
  • Run three pilot interviews on your hardest segment, not your easiest, and read the raw transcripts rather than the synthesis.
  • Mark every question that suggests its own answer. Adversarial reading catches leading phrasing that looks neutral to its author.
  • Check the first sentence of each section for an assumption the participant has not yet confirmed.
  • Have a native speaker read it in every language you will field, at the register your participants actually use, since segment vocabulary rarely survives literal translation.

The pilot transcripts are where the real signal is. Look for questions that produced identical answers across all three participants, which usually means the question told them what to say. Look also for questions everyone misunderstood in the same way, which points at vocabulary rather than logic.

Budget a full rewrite after the pilot rather than a tidy-up. Guides written for AI moderation tend to fail structurally rather than cosmetically, because the problem is usually an assumption baked into the order or the probing map rather than a badly worded sentence. Teams that plan for one revision cycle ship better studies than teams that treat the pilot as a formality.

What the Guide Cannot Fix

Being honest about the ceiling keeps the effort proportionate.

  • The unanticipated finding stays out of reach. No amount of pre-written probing substitutes for a moderator noticing something nobody planned for. For genuine discovery these tools are not yet suited to semistructured, in-depth work, and that boundary is still moving.
  • Nonverbal signals are thin. The interviewer presents no face, and outside video modes it reads none either.
  • A perfect guide on the wrong sample is still wrong. Mode and recruitment decide who answers.
  • Consistency amplifies whatever you wrote, including the flaws, identically across every participant.
  • Consent and welfare are not guide problems but they are your problems. The ESOMAR code and guidelines require that disclosure be genuinely understood, and the Insights Association and AAPOR publish complementary standards.

Platform defaults make guide decisions on your behalf unless you override them, and two are worth setting deliberately. Guides for individual interviews work best at roughly six to eight primary questions, and question order changes the answers that follow it, which under a fixed script applies identically to everyone rather than varying between sessions.

Frequently Asked Questions

How is a guide for an AI moderator different from a normal one?
That depends on how much latitude the moderator has. On a rigid system it carries every judgment a human interviewer would make live. A person skips questions already answered, rephrases confusing ones and follows unexpected answers. An AI moderator does none of that, so skip logic, phrasing and probing depth all have to be decided in the document before fieldwork starts.
How many questions should an AI-moderated interview have?
Roughly 8 to 12 substantive questions for a 20 to 25 minute session, with probing budget included, since a deeply probed question effectively costs two or three. Design to the length rather than trimming afterward, because AI-moderated sessions manage time poorly and will not self-correct if you overrun.
How much probing should each question get?
Deep probing belongs on the two or three questions your decision actually rests on, one follow-up where a one-word answer would be useless, and none on warm-ups or screeners. Current tools let you set probing per question. Depth decides where your findings will be rich, so distribute it by decision rather than by hypothesis.
How do you avoid leading questions in an AI-moderated guide?
Read the guide adversarially and mark every question that presumes its own answer, such as asking what frustrated someone before establishing that anything did. With a human moderator a leading question gets softened in the room. With an AI it is asked verbatim to every participant, so it becomes a systematic distortion.
Should the guide change for WhatsApp or phone interviews?
Yes. Messaging interviews need shorter questions, no long preambles and tolerance for voice-note answers. Phone interviews cannot use visual stimulus, so anything shown must be described in words that work read aloud. Asynchronous formats need each question to stand alone across a gap of hours.
How do you test a discussion guide before fielding?
Read it aloud end to end, have someone answer it deliberately badly to see whether the probing rules recover anything, then run three pilot interviews on your hardest segment and read the raw transcripts. Questions that produced identical answers across all three usually told participants what to say.

About the Author

Sreenadh Narayanan is the founder of Alchemic, an AI-powered consumer research platform used for ad testing, concept testing and brand tracking. He writes Alchemic's guides on qualitative research and research methods, covering interview design, sample sizes and how teams turn customer conversations into decisions.