Last updated: 18 September 2026
Quick Answer: New product research works best as one study per development stage, each sized to the decision it clears. Six stages cover a launch: idea screening, concept, product use, price and pack, a pre-launch forecast, and the in-market read. Concept tests usually need 200 or more respondents; a product-use test needs far fewer.
New product research is not one study. It is a short sequence, each study retiring one uncertainty and returning a result clear enough to stop the program if it comes back wrong. One of the few studies to test development practice against product performance, an examination of ten financial services providers in Nigeria, traced underperformance to overestimated market size and inadequate research rather than bad ideas.
The expensive mistake is not skipping research. It is commissioning a study whose result cannot change the decision in front of you. A concept test returning 62% top-two-box purchase intent tells you nothing unless somebody agreed beforehand which number would have stopped it. Stated intent is not purchase: in the 2,000-household automobile survey Baohong Sun and Vicki Morwitz analyze in the International Journal of Research in Marketing, 9.9% said they would buy within a year and 5.0% did.
The generic research process does not change for a new product: defining the objective, choosing a method, recruiting and analyzing works the same whatever you study. What changes is that every stage measures a proxy, and the proxies weaken the further they sit from a purchase.
What Four Uncertainties Does Product Research Have to Retire?
Four, and they have to fall in order: whether the problem is real, whether your solution fits it, whether anyone will pay, and how it should be positioned. A study aimed at an uncertainty the previous stage never settled returns a confident number about the wrong question.
- Does the problem exist? Open-ended conversation with people who have it, before you describe anything.
- Does your solution fit it? Concept exposure, comprehension checked before any rating.
- Will they pay? Price sensitivity against a settled proposition, never inside a design test.
- How should it be positioned? Message and pack testing, once the first three close.
Federal guidance follows the same hierarchy. The US Small Business Administration's market research and competitive analysis guide puts demand and market size first and pricing last. A price answer collected before the demand answer is a preference among strangers.
Which Study Retires Which Uncertainty?
Ordered by development stage, earliest first, which is also the order the uncertainties settle in.
| Stage | Uncertainty it retires | Evidence that clears it | Typical sample | Result that should stop the program |
|---|---|---|---|---|
| Idea screening | Is the problem real and worth solving? | Depth interviews on current behavior, plus a light quant read | 9 to 24 interviews per audience | No idea separates on unaided relevance |
| Concept testing | Does this solution fit the problem, and for whom? | Monadic exposure, comprehension checked before rating | 200 or more, commonly 100 to 400 | Respondents cannot restate the benefit |
| Product and prototype use | Does the thing deliver the promise? | In-home use, central location testing, live prototype sessions | 8 to 15 per persona for interface work; more for sensory | Repeat-use intent falls below first-use intent |
| Price and pack | Will they pay, and does the pack carry the promise? | Price sensitivity first, then shelf and visual testing at the settled price | 50 to 200 for packaging | Preference reverses once price is shown |
| Pre-launch forecast | How much of this will actually move? | Trial and repeat from adjusted intent, with assumptions stated | 200 to 500 per market | The forecast only works at distribution you cannot buy |
| In-market read | Did any of the above hold? | Early sales, returns, reviews, a tracking wave against baseline | 200 or more per wave | Trial lands and repeat does not |
Those interview norms are measured, not habitual. Hennink, Kaiser and Marconi, running 25 in-depth interviews for Qualitative Health Research, found 91% of codes surfaced by the ninth, but 16 to 24 before the meanings behind them were understood.
The last column earns its place. Agreeing the stopping result before fielding separates research from reassurance. A stop rule written afterward is not one: the number is known by then.
Stage order, and what contaminates a run taken out of turn, is covered in the launch research sequence.
How Do You Screen Ideas Without Killing the Good One?
Screen on the problem, not on the idea. Ask what people currently do, what it costs them and what they have already tried, then watch which of your ideas anyone raises unprompted. Scoring concepts on appeal this early promotes the most familiar option, because familiarity is the easiest thing for a stranger to rate.
The saturation numbers above are the screening trap in one line. Nine interviews gave Hennink and colleagues the range of issues; understanding those issues took 16 to 24. Nine interviews buy you the vocabulary. Choosing which idea to develop on nine is choosing on vocabulary rather than on need, which is how teams kill the unfamiliar idea and fund the obvious one.
Those figures came from a homogeneous sample of 25 participants at a single clinic, and the authors cautioned against generalizing them, so treat them as a floor rather than a rule. A study spanning three countries and four income bands needs materially more.
Secondary data does real work here and costs nothing. County Business Patterns reports establishments, employment and payroll across nearly 1,000 industries down to ZIP code level, and Census Business Builder profiles a geography before you field anything. Market saturation is a desk question. Do not spend fieldwork on it.
When Does the Product Itself Have to Be in the Room?
As soon as the promise depends on execution rather than on the idea. A concept test measures a description, and a description cannot disappoint anyone. A product test measures the thing, and the gap between the two is where launches quietly fail.
For interface and digital products the useful observation is behavioral rather than stated: hesitation, a mis-click, a back-out, probed while the participant is still in the flow, not in a debrief afterward. The 8 to 15 per persona figure in the table is a working floor, not a guarantee. A large-scale evaluation by Spool and Schroeder, reported in a review of usability testing practice by J.M. Christian Bastien, ran 49 participants across four shopping sites. The first five uncovered 35% of the usability problems, and the serious ones that actually blocked a purchase did not surface until the 13th and 15th participant. That study is worth the space because its task was buying something, which is the condition a product test is trying to reproduce.
For physical goods the constraint is time, not sample. If a shampoo takes three weeks of use to judge honestly, it takes three weeks, and no fieldwork method compresses a wear period. A day-one read measures packaging and first impression, not the product.
How Do You Test Price and Pack Without Confounding Them?
Separately, and price first. A pack test run with price visible is measuring value, and its ranking of the designs is not recoverable afterward, because design preference and perceived value move together in the response.
The workable sequence is price sensitivity questioning against the settled proposition, then shelf and visual testing at the price that emerged. Finding out what customers will actually pay is its own method, and it belongs before the design work. Packaging studies run 50 to 200 respondents, enough to separate designs on standout and communication, never enough to read an elasticity.
Stated prices are not automatically inflated. A choice experiment on the carbon footprint of mandarin oranges in PLOS ONE compared 104 participants spending real money against hypothetical samples of 212 in the lab and two online groups of 500, and measured willingness to pay of 0.53 yen per gram of emissions reduced against 0.52 to 0.58 hypothetically, with no significant gap in any of the six tests. The wider picture is context-dependent: the synthesis of stated-choice evidence by Haghani and colleagues reports negligible hypothetical bias in health choice experiments and significant bias in consumer and transport settings.
The risk in a price study is rarely that respondents lie about money. It is that the design let something else move at once. A reversal is the tell: if design A leads without price and design B leads once price appears, you have learned nothing about design and a lot about an unsettled proposition.
What Can a Pre-Launch Forecast Predict, and When Do You Find Out?
It predicts trial reasonably and repeat poorly, on assumptions that are business decisions rather than research findings. A volumetric model converts adjusted intent into a trial estimate, then multiplies it by awareness and distribution. Research supplies the first number. The other two are things the company chooses to buy. How far stated intent overstates the purchase that follows is set out in what purchase intent scores actually predict.
The adjustment is not optional. The personal computer sample in the same Sun and Morwitz paper makes the point on a second category: 15.3% of respondents said they would buy a computer within a year and 5.7% did. Different product, different decade, same direction.
This is where market research for product launch is most often oversold. The Nigerian study cited above traced product underperformance to overestimated market size, and that kind of overestimation rarely comes from the intent data. It comes from the assumption layer wrapped around it, where the plan's distribution is not the distribution the sales team lands.
Two disciplines keep a forecast honest.
- State every assumption as a line the business owns, so a miss lands on the plan rather than on the research.
- Set the in-market checkpoint before launch. Trial and repeat separate within two or three purchase cycles, the cheapest diagnosis available. Strong trial with weak repeat points at the product; weak trial with strong repeat points at awareness, distribution or price.
Who Does Your Sample Miss, and What Does That Cost the Launch?
Everyone who cannot or will not take a browser survey, which in most consumer categories is a large and non-random share of buyers. The error runs one way, and it flatters new products.
Coverage is the problem. The ITU's Facts and Figures 2025 records almost three-quarters of the world online and 2.2 billion still offline, most in low and middle income countries. A browser-video study in a growth market samples the connected top slice: India's statistics ministry put internet use among people 15 and over at 85.5% for urban men against 57.6% for rural women in early 2025. That frame fits a premium skincare launch, not a detergent. Coverage is also not adoption: the ITU report above puts 5G within range of more than half the world's population, and at more than a third of all mobile broadband subscriptions.
The gap is not only a poor-country problem: Pew Research Center finds 16% of US adults are smartphone-only, owning a phone but no home broadband: a browser-video study loses them first.
Channel changes who answers. Alchemic runs interviews natively inside WhatsApp with no link and no app, and AI phone interviews to any working number, feature phones included. It publishes 57+ languages including Spanish, Arabic and Mandarin, with managed fieldwork or bring your own panel across fourteen markets including the USA and the UK.
Where a Browser Panel Already Reaches Everyone
Reach is not always the binding constraint. For a B2B product sold to an English-speaking, permanently online US buyer, a browser panel reaches the whole universe, costs less and fields faster; multi-channel fieldwork there solves a problem that does not exist. Source still predicts quality: in a November 2024 opt-in survey of 11,114 US adults, Pew Research Center found 18% said yes to at least one trap question designed to catch insincere answers. Disclose the sample source and recruitment method, as the AAPOR standards ask, so the result stays auditable.
What Product Research Cannot Settle Before Launch
Execution, competitive response and category creation. No stage above tests whether the sales team lands the listing, whether a rival cuts price in week three, or whether a retailer gives the facings the forecast assumed.
Genuinely new categories are the sharpest limit. Respondents evaluate against a reference they already hold, so a product with no category gets scored against the nearest familiar thing. Discovery work with human moderators still outperforms every faster option here: the value is in noticing the question nobody thought to ask, and a moderator who can abandon the guide mid-session is the instrument for that. A normative concept benchmark from a syndicated database is the other thing no interview program replaces.
Guide rigidity is a real risk, and how much you carry depends on the system. A rigid instrument makes the guide the entire study, while a more dynamic conversation keeps it a starting point. Once it has the client's brief, Alchemic's researchers tailor the discussion guide to handle these risks before fielding, rather than leaving that work to the buyer, and how far a moderator may depart from that guide mid-interview is worth asking any vendor separately. Across every route the ESOMAR code and guidelines govern consent and participant welfare, and the Insights Association publishes complementary guidance.

