Last updated: 19 August 2026
A launch research sequence is the order in which a brand tests an idea, a concept, a product, a pack and its creative. In FMCG and wider consumer goods, each stage answers exactly one question well. Running them out of order is the most common and most expensive research mistake a brand team makes.
The sequencing problem has become more acute rather than less as research got faster. When a concept test took six weeks, the calendar itself enforced order.
Now that AI moderation has compressed qualitative fieldwork from weeks to days, teams can and do run stages in parallel. That feels efficient. It quietly destroys the logic that made each stage interpretable.
The respondent base has moved just as fast. DataReportal's Digital 2026 report counts more than 6 billion people online, while the ITU records 2.2 billion still offline, which is the gap that decides whose reaction a launch study captures.
This article sets out the sequence, what each stage can and cannot answer, and where compression genuinely helps.
Why Does Sequence Matter More Than Method?
Because every research stage assumes the previous one is settled. A concept test assumes the need is real. A pack test assumes the concept is agreed. A creative test assumes the pack and proposition are fixed.
When an assumption underneath a stage is still open, respondents answer the open question instead of the one you asked, and the result reads as a clean finding rather than as contamination.
This shows up in a recognizable way. A pack test where respondents keep mentioning price is not a pack test. It is a value-proposition test wearing a pack test's clothes, and its ranking of pack designs is close to meaningless, because design preference and perceived value are confounded.
The discipline is straightforward: settle the question upstream before you spend on the question downstream.
The Sequence, Stage by Stage
| Stage | The one question it answers | Method that fits | What it cannot tell you | Typical decision |
|---|---|---|---|---|
| Need and opportunity | Is there an unmet need worth serving? | Exploratory qualitative, human-moderated | Whether your specific idea wins | Whether to enter at all |
| Idea screening | Which of these ideas has any pull? | Lightweight quant with open-ends | How to execute the winner | Which two or three to develop |
| Concept testing | Does the proposition land, and with whom? | Monadic qualitative and quantitative | Whether the pack or ad works | Which concept to build |
| Product and sensory | Does the product deliver the promise? | In-home use, central location | Whether people will buy it | Reformulate or proceed |
| Packaging | Does the pack communicate and find itself on shelf? | Visual testing, shelf simulation | Whether the price is right | Which design to print |
| Creative and ad testing | Does the execution communicate the settled idea? | Ad testing, qualitative reaction | Whether the concept is good | Which cut to run |
| Pre-launch tracking baseline | Where does the brand start? | Brand tracking wave zero | Post-launch causation | The benchmark for everything after |
Two notes on reading this table. The one-question column is a constraint, not a summary: each stage genuinely answers that question and answers adjacent ones badly. And the baseline is a stage, not an afterthought. Teams that skip wave zero cannot attribute anything afterward, which is the single most common reason launch post-mortems are inconclusive.
Where Does AI Moderation Actually Help?
At the stages where you already know what to ask. AI-moderated interviews are strongest there and weakest everywhere else. The Nielsen Norman Group's January 2026 testing found the two platforms it tried worked best for structured input at scale, naming product feedback and screening as strong fits, while sticking to the script rather than chasing unexpected insight.
Mapped onto a launch:
- Idea screening: Best for AI moderation. High volume, known questions, open-ends that would otherwise go uncoded.
- Concept testing: Strengths: monadic designs at a sample size that was previously unaffordable, with real reasons-why rather than scale ratings. Limitations: a weak question still yields weak answers; a rigid moderator will not rescue it, so the guide has to be sound whether the buyer writes it or the vendor's researchers do.
- Creative and ad testing: Strengths: fast reaction across many cuts and segments, which makes creative testing timing far more forgiving than it used to be. Limitations: no read on nonverbal response.
- Need and opportunity: Best for human moderators. That testing is explicit that these tools are not yet suited to semistructured, in-depth discovery interviews, and this is the stage that is purest discovery.
The pattern is consistent. AI moderation is strongest in the middle of the funnel, where the question is already known and the advantage comes from volume and speed. It is weakest at the mouth of the funnel, where the advantage comes from spotting the question nobody had thought to ask.
Who Is in Your Launch Sample?
For consumer goods this question is unusually consequential. Volume in most categories sits with mass-market consumers, not the affluent urban segment that online research reaches most easily.
Pew Research Center's mobile technology fact sheet reports 16 percent of US adults as smartphone-only internet users, rising to 34 percent in households under $30,000 a year against 4 percent above $100,000. A browser-video study therefore samples unevenly across exactly the income gradient a mass-market launch depends on, and the ITU's Facts and Figures 2025 shows the same gradient running far wider across other markets, with 2.2 billion people still offline.
For a premium skincare launch that may be the right frame. For a detergent, a value snack, a haircare line or an entry-price personal care product it is the wrong one. The error runs in a predictable direction. The sample over-indexes on consumers who are less price-sensitive, more brand-aware and better served by broadband, which flatters concepts that would struggle in the actual market.
Alchemic addresses this by running interviews natively inside WhatsApp, with no link and no app to install. It also runs AI phone interviews to any working number, including feature phones. The same selection effect is examined in sample validity and who you miss.
Fieldwork is managed or bring your own, with moderation across 57+ languages and delivery across fourteen markets, spanning the USA and the UK as well as South and Southeast Asia, the Gulf and Africa. Brands including Unilever, Mars, Dr. Reddy's, Sleepwell and CaratLane appear on its client roster. Whether that reach matters depends on the category: the more mass-market the product, the more the sampling frame decides whether the concept test predicted anything. Where that sample comes from in the first place is covered in survey panels and where respondents come from.
What Can You Compress, and What Cannot Be Compressed?
Speed is a genuine gain, but it is not evenly available across stages.
Compressible: idea screening, concept iteration, creative reaction, and the open-ended coding that used to bottleneck every stage.
Running eight concept variants instead of three is a real methodological improvement, not just a faster one, because it lets you find the edges of a proposition rather than picking among three pre-selected guesses. Per variant the sample can stay modest: qualitative work finds code saturation arriving at around nine interviews even where fuller meaning takes longer to settle, so breadth across concepts usually buys more than depth within any one of them.
Not compressible:
- In-home use tests. If the product takes two weeks to evaluate honestly, it takes two weeks. No moderation change touches this.
- Sensory and taste work. The product has to be in the respondent's hands.
- Physical shelf context. Screen-based shelf simulation is useful and is not the same as a store.
- Discovery. Compressing the front end is how teams end up testing well-executed answers to the wrong question.
- Category learning cycles. Seasonal categories have real calendars.
The honest framing is that AI moderation compresses the fieldwork and analysis portion of a stage. That is often most of the elapsed time. It is never all of it.
Why Is the Baseline the Riskiest Stage to Cut?
The baseline stage deserves a closer look. It is the one most often cut when a launch calendar tightens, and the one whose absence is least recoverable.
A pre-launch wave costs little and takes little time, and it is the only artifact that makes every subsequent measurement interpretable. Awareness rising from an unknown starting point is not a result. Cutting it converts a year of post-launch tracking into numbers with nothing to compare against. Treat it as a fixed cost of launching rather than a line item competing with the others.
A second underrated compression is running the qualitative and quantitative halves of a stage together rather than sequentially. Concept testing historically meant a qualitative round followed by a quantitative confirmation. A single program can now carry both, provided the open-ended component is genuinely designed rather than bolted onto a survey as a free-text box.
Where This Sequence Breaks Down
The stage model is a default, not a law, and several common situations legitimately depart from it.
- Line extensions can skip the front end. If the need and the brand are established, starting at concept or even pack is reasonable.
- Genuinely new categories may need to loop. Concept and product findings can reopen the need question, and that is a finding rather than a failure.
- Parallel is sometimes right. Running pack and creative together is defensible when both are downstream of a genuinely settled proposition and the timeline is fixed by a retailer window. Know that you are doing it and why.
- Retailer and regulatory calendars override everything. A listing window or a labeling deadline is a real constraint. A program compressed to meet one is a considered tradeoff rather than a mistake, provided the team records which stage was shortened.
- Sample quality beats stage discipline. A perfectly sequenced program run on the wrong 8% of the market is worse than a slightly out-of-order one run on the right people.
- Some stages need human moderation regardless of budget. For discovery and for emotionally complex categories, independent testing found human interviewers still outperform.
Where Do These Stage Rules Come From?
Professional standards apply throughout. The ESOMAR code and guidelines govern consent and participant welfare at every stage, and the Insights Association and AAPOR publish complementary guidance.
The sequencing argument is not new to consumer goods. Robert Cooper's Stage-Gate framework puts the leading cause of new-product failure on a lack of understanding of customers, their problems and their purchasing criteria, which is precisely what an out-of-order program fails to build. Work on product development practice and failure finds the two compound: when poor execution follows inadequate development practice, the likelihood of failure rises.

