Last updated: 23 September 2026
Quick Answer: To run customer discovery interviews at scale, saturate each buyer segment, recruit verified buyers, and run one consistent guide in parallel. Empirical tests put saturation at 9 to 17 interviews per homogeneous group, so 100 interviews should mean roughly eight segments, not one crowd. Code themes while fieldwork runs, against a fixed codebook.
Steve Blank told founders to get out of the building, and most teams still read that as a call count. When the business outgrows ten calls, with three buyer types, two markets and a board wanting more than anecdotes, the reflex is to book 100. With no segment plan, 100 calls produce a louder version of the same ten.
Scale in discovery is a segment count, not an interview count. Each distinct buyer group has to reach saturation on its own, and a large pile of interviews drawn from whoever answered first saturates nobody except the loudest group in it.
Which Best Practices Hold When Discovery Interviews Scale?
The craft rules do not change at volume; their failure cost does. Ask about specific past behavior, never about your idea. A leading question asked once is a bad interview. Asked 200 times, it is a biased dataset that looks like evidence.
Rob Fitzpatrick's book The Mom Test is built on that rule, and the basics for a founder running a dozen calls are covered in our guide to market research for small businesses and startups. Three practices matter more once the count climbs:
- One guide, locked before fielding. Every interviewer, human or AI, works from the same opening questions and the same probes, so differences in answers come from respondents, not from who asked.
- Hypotheses written as segment claims. "Finance leads at 50 to 200 person firms reconcile invoices by hand" can be confirmed or killed. "People hate invoicing" cannot.
- A stopping rule per segment. Decide in advance what counts as enough, or the program stops when the budget does.
How Many Discovery Interviews Does a Decision Need?
Plan 9 to 17 interviews per homogeneous segment, then multiply. A 2022 systematic review by Monique Hennink and Bonnie Kaiser of 23 studies that tested saturation found that studies using empirical data reached saturation within 9 to 17 interviews or 4 to 8 focus group discussions, mainly with homogeneous populations and narrow objectives. The outliers, including multi-country research, needed larger samples, which is exactly the at-scale case, so treat the per-segment count as a floor.
Guest, Namey and Chen showed how to call saturation during fieldwork rather than guess it beforehand. Their 2020 method in PLOS ONE uses a base of four interviews, runs of two, and a new-information threshold of 5 percent or less. Across three bootstrapped datasets, six to seven interviews captured the majority of themes in a homogeneous sample, and the most varied dataset needed more.
Planning arithmetic, sorted alphabetically by decision type. These are budgeting assumptions built on the ranges above, not findings:
| Decision type | Segments to saturate | Interviews per segment | Planned total |
|---|---|---|---|
| Entering a second country | 3 per market, 2 markets | 12 to 17 | 72 to 102 |
| Picking an ideal customer profile | 4 to 6 candidate segments | 12 | 48 to 72 |
| Pricing and packaging discovery | 3 buyer roles | 12 to 15 | 36 to 45 |
| Validating one problem in one segment | 1 | 12 to 17 | 12 to 17 |
How Do You Recruit Real Buyers at Volume?
Recruit from verified buyer lists and screen on behavior, because an open paid invitation at scale attracts people who want the incentive, not people with the problem. In a 2024 study in BMC Health Services Research, Sharma and colleagues advertised on Facebook for carers of Australian children, offering 50 Australian dollars. All 254 people who expressed interest were judged likely imposters, posing as multiple individuals from IP addresses across Nigeria, Australia and the United States.
It is one health-services study, cited because it documents the failure end to end. The red flags transfer directly to buyer recruitment: bursts of sign-ups at odd hours, unlikely demographics, short or vague interviews, a preference for camera-off participation, and fixation on payment. The countermeasure is a screener that asks about past purchase behavior an imposter cannot guess, covered in our guide to survey screening questions.
Reach is the second problem. A browser link reaches people at a desk who open research email; shop owners, field technicians and buyers outside English-speaking metros need channels they already use. Alchemic runs interviews natively inside WhatsApp, with no link and no app, and places outbound AI phone calls, with managed fieldwork or bring your own recruitment across 14 markets including the USA and the UK. Sourcing trade-offs are covered in our note on where good respondents come from.
What Changes When an AI Moderator Runs Discovery?
An AI moderator removes the calendar as the constraint: hundreds of interviews can run at once, in the respondent's language, on one guide. Someone still has to design the guide, check the recruit and read the output. How much latitude a system has to leave the guide and chase an unexpected answer varies by tool, so test that on a pilot before fielding at volume.
ESOMAR's 20 Questions to Help Buyers of AI-Based Services is the vetting list to use. Its five sections cover:
- The supplier's company profile and credentials.
- Whether the AI is explainable and fit for purpose.
- Whether it is trustworthy, ethical and transparent.
- How humans oversee it.
- Data governance.
For discovery, oversight matters most: who reads transcripts during fieldwork and can stop a guide producing noise.
Once it has the client's brief, Alchemic designs and tailors the discussion guide before fielding rather than leaving that work to the buyer, publishes 57+ languages including Hindi, Tamil and Telugu, and runs a 200-interview qualitative study from brief to live dashboard in 3 days, 5 to 7 for complex studies. What a guide needs when an AI runs it is set out in our piece on the discussion guide for an AI moderator.
How Do You Turn Hundreds of Transcripts Into a Decision?
Code while fieldwork runs, per segment, against a codebook fixed after the base interviews, and report behaviors counted by segment rather than opinions averaged across everyone. A theme that appears in 11 of 12 finance leads and 1 of 12 operations leads is two findings, and pooling them into "50 percent" destroys both.
Tie every claim to a verbatim. Alchemic auto-codes themes while fieldwork runs and links every claim to the respondent's verbatim or voice clip, and its knowledge base carries across waves rather than restarting each study. For teams coding by hand, our comparison of qualitative data analysis software covers the tooling.
The Copyable Run Sheet
Copy this into the study plan and fill every line before the first interview:
- Decision: the one choice this program informs, and the date it gets made.
- Segments: each buyer group written as a testable claim, with a target of 12 per segment.
- Stopping rule: base of 4, runs of 2, stop at 5 percent or less new information, never before interview 9.
- Screener: two behavioral questions about a real past purchase, plus one trap answer.
- Channels: where each segment actually answers (email, WhatsApp, phone, in-product).
- Guide: five past-behavior questions, zero questions about your idea, probes written in.
- Pilot: 4 to 6 interviews, read in full, guide revised once, then locked.
- Codebook: frozen after the base interviews; new codes logged with the interview they came from.
- Readout: themes by segment, a verbatim per claim, and the hypotheses killed.
Where Does Discovery at Scale Fall Short?
Volume does not rescue a wrong question. Three cases where another approach is the better choice:
- One segment, one problem, early stage. A founder making ten calls personally learns faster than any program. Scale adds cost before signal.
- Behavior people cannot report. Workarounds, shelf choices and in-app hesitation are better observed than recalled, through analytics, field observation or a usability session.
- Proving demand. Discovery finds problems; it does not size or prove them. A survey sizes them and a live test proves them.
Multi-country programs carry the heaviest saturation burden. Budget the per-market floor rather than splitting one market's sample across three.

