Last updated: 4 September 2026
A purchase intent score is the share of respondents to a purchase intent survey who say they would buy, conventionally the top two points of a five-point scale. It predicts the ranking of one concept against another reasonably well, and the size of actual demand badly. Two jobs, one number. Almost every argument about purchase intent is really an argument about which of them the score is being asked to do.
The scale of the second failure is old and well documented. Analyzing United States Bureau of the Census data on new car purchases, Theil and Kosobud found that 70 percent of automobile purchases came from people who had stated no buying intention at all. Fewer than 40 percent of those who did state an intention actually bought a car. Their result is reported in the Marketing Bulletin review of purchase prediction by Day, Gan, Gendall and Esslemont.
Both halves matter, and most treatments mention only one. Intenders overstate. Non-intenders also buy, in large numbers, and no deflation factor recovers them. Purchase intent is a filter with false positives on one side and invisible buyers on the other, which is why it behaves like a comparative instrument and misbehaves like a forecast.
How the Score Is Built: Scales, Boxes and Conventions
The standard instrument is a five-point verbal scale running from definitely would not buy through might or might not to definitely would buy. Top-two-box, usually written T2B, sums the two most positive points into one percentage; top box reports only the most positive point. The convention survives because it is legible, and what it discards is composition, which carries commercial meaning of its own.
- A 40 percent T2B built from 30 percent definitely and 10 percent probably is a different commercial proposition from the same 40 percent built the other way around. Reporting T2B alone erases the difference.
- Verbal points are not interpreted consistently between respondents. Two people can both select probably would buy while meaning very different probabilities, and the analysis treats them as identical.
- The scale cannot describe non-intenders at all. Everyone below the top two points collapses into a single discarded residual, and that residual is where a large share of real purchases originates.
The main alternative addresses that last problem. Juster proposed an eleven-point purchase probability scale, arguing that verbal intentions are disguised probability statements and probabilities should be collected directly. Each point carries an explicit odds equivalent, from no chance at zero through five in ten to practically certain at ten.
Because it yields a mean probability rather than a proportion of intenders, it assigns non-zero probability to people a verbal scale writes off. Day and colleagues report that in Juster's own study, purchase probabilities explained twice as much of the variance in actual purchase rates as buying intentions data.
Report top box, second box and T2B together. It costs nothing and stops a composition change reading as a demand change.
Why Does Stated Intent Overstate Real Demand?
Because saying yes in a survey is free, and buying is not. Three mechanisms drive the inflation, and they compound.
Hypothetical bias. Where the same elicitation runs hypothetically and for real money, the hypothetical answer is systematically higher. Murphy, Allen, Stevens and Weatherhead's 2005 meta-analysis in Environmental and Resource Economics examined 28 stated preference studies that used the same mechanism for both, generating 83 observations. The median ratio of hypothetical to actual value was 1.35, with severe positive skew. The skew is the important part: the typical study inflates modestly, a minority enormously.
Social desirability. Respondents adjust answers toward what looks good, and the effect is measurable and partly fixable. A systematic review of social desirability bias reduction methods covering 121 experiments across 79 papers found bias significantly reduced in 55 percent of experiments, with face-saving designs succeeding in all 18 that used them and mode of administration in half. How a question is framed and delivered is not cosmetic.
The intention-behavior gap itself. This is the part most easily misread, because the correlational evidence looks strong. Sheeran and Webb's review of the intention-behavior gap reports a sample-weighted average correlation of r+ = 0.53 between intention and later behavior across ten meta-analyses covering 422 studies.
Yet Webb and Sheeran's meta-analysis of experiments that actually manipulated intention found a medium-to-large change in intentions produced only a small-to-medium change in behavior, d+ = 0.36. The people responsible for the gap are the ones who intend and then do not act, described in that literature as inclined abstainers.
Those two numbers are the argument. Intent correlates with behavior, so it ranks well; moving intent moves behavior far less, so it forecasts badly.
What Deflation Factors Do, and What They Hide
A deflation factor is a multiplier applied to a T2B score to convert it into an expected trial or purchase rate. It is a formalized admission that the raw number was never a forecast.
Applied honestly it is calibration: the multiplier comes from the same category, wording, sample frame and horizon, derived from cases where both the score and the outcome are known. Applied as a house constant it is decoration, because a multiplier reused across categories imports assumptions nobody can check when the derivation is proprietary.
The published work names the conditions that actually govern the relationship. Morwitz, Steckel and Gupta's analysis in the International Journal of Forecasting, volume 23, 2007, identified six.
Intentions track purchases more closely for existing products than new ones, for durables than non-durables, and over short horizons than long ones. The relationship also tightens when respondents rate specific brands or models rather than a product category. Same when the outcome is trial rather than total market sales, and when intentions are collected comparatively rather than one concept at a time.
Read that against a typical concept test and the tension is obvious. Concept tests usually run on new products, often at category level, frequently monadic, against a distant launch. Those are the conditions under which the relationship is weakest. Monadic designs remain right for many studies, because rating a concept in isolation removes a comparison the shopper would never make; the design choice protecting against one bias weakens the predictive relationship at the same time.
Three questions make a deflation factor auditable: what category and horizon it came from, how many calibration cases sit behind it, and what wording those cases used. A supplier who cannot answer is offering a number, not a calibration.
Which Intent Question Should You Ask?
Buying intent research has no single best instrument. Each form trades reach against precision, and the choice follows the decision the number has to support.
| Question form | What it asks | Predicts well | Known weakness |
|---|---|---|---|
| Five-point verbal, T2B | Would you buy, on definitely-to-definitely-not | Relative ranking between concepts tested the same way | Inflates; hides composition; cannot describe non-intenders |
| Top box only | The definitely-would-buy share alone | Strength of the committed core | Discards softer demand that often converts |
| Juster 11-point probability | Odds out of ten that you buy in a period | Aggregate purchase rates, especially durables | Longer to administer; needs a defined time window |
| Yes / no or three-point | A binary or coarse intention | Fast screening and very large samples | Least information per respondent; crude |
| Choice among alternatives | Which of these would you buy | Share-style questions and competitive context | Not an absolute demand estimate |
| Behavioral proxy | A pre-order, deposit or real-money choice | Actual conversion, because it is a purchase | Expensive; needs a real offer; late in the process |
Best for most concept work: the five-point scale, reported as three numbers rather than one, read against a benchmark from the same category. Best where the decision is expensive: a behavioral proxy. A real deposit or a live pre-order page beats every survey instrument here, because it removes the hypothetical entirely. Survey intent is what you use when a behavioral test is not yet possible, not a substitute for one that is.
How Much Does Question Wording Move the Number?
More than most teams assume, and often more than the difference between the concepts being tested.
Pew Research Center's work on writing survey questions shows the size of the effect. In a January 2003 experiment, support for military action in Iraq ran at 68 percent in favor against 25 percent opposed, then fell to 43 against 48 once the question added that it might mean thousands of US casualties. One clause moved the result 25 points. Question order shifted support for another item by eight.
Three decisions matter most for the purchase intent question specifically.
Whether comprehension is checked before the rating. A respondent who misread the concept still produces a score, and it enters the average indistinguishable from an informed one. Asking what the product is and who it is for, before asking whether they would buy, separates a comprehension failure from a rejection.
Whether price is present. Intent without price measures appeal; intent with price measures something closer to demand. They are not comparable and should never be trended against each other.
Monadic or sequential monadic. Rating in isolation avoids artificial comparison; rating several in sequence produces the comparative context Morwitz and colleagues associate with stronger prediction, at the cost of order effects a rotation absorbs.
Does Asking the Question Change the Answer?
Yes, and by enough to matter. Measurement is not passive here. Morwitz, Johnson and Schmittlein showed in the Journal of Consumer Research in 1993 that merely asking about purchase intent raises the subsequent purchase rate, later called the mere-measurement effect. Chandon, Morwitz and Reinartz found in the Journal of Marketing in 2005 that surveying inflates the apparent link between intentions and behavior, an effect known as self-generated validity.
Some of the correlation that makes intent look predictive is created by the survey. The frame underwriting the question is the theory of planned behavior, in which intention is formed by attitude, subjective norm and perceived behavioral control, and it organizes current work predicting purchase behavior.
Which Buyers Never Reach Your Intent Score?
Sample composition moves a purchase intent number as much as any scale decision, and it is the least reported part of a concept test. Research recruited only through browser-based panels under-covers parts of the market a mass brand actually sells to.
Pew Research Center's mobile technology fact sheet reports 16 percent of US adults as smartphone-only internet users. That rises to 34 percent among adults in households under 30,000 dollars a year, against 4 percent above 100,000, and the ITU's connectivity statistics show a steeper gradient across markets.
Those respondents are not missing at random. They skew lower income and less urban, and for a value proposition their intent is the number that should decide the launch. Language does the same more quietly: a concept rated in a second language is rated through a translation the respondent performed, and comprehension failures surface as lower intent indistinguishable from rejection.
Alchemic's concept testing service checks comprehension before it asks for a rating, supports monadic and sequential monadic designs, and probes the reason behind each score in the same interview. Sample norms are 200 or more, with a working range of 100 to 400, and ad testing runs at 200 to 500 per market.
Fieldwork reaches respondents natively inside WhatsApp, with no link and no app, and by outbound phone call as well as by browser, with the respondent choosing. Alchemic publishes 57+ languages including Hindi, Tamil, Telugu, Bangla, Arabic and Indonesian, and recruitment runs as managed fieldwork across 14 markets including the USA and UK, or bring your own panel.
When Intent Data Should Not Decide a Launch
Four situations where the number should inform a decision without making it.
Genuinely new categories. Intent is weakest exactly where products have no reference point, which is where launch risk is highest, and concept tests routinely run a year or more ahead of shelf. A score for a product nobody has encountered measures reaction to a description.
Absent price, distribution or competition. A concept tested without a price, in a category the respondent will meet on a crowded shelf, has been tested under conditions that will never occur.
Small cells. Read by market, age band and variant, a generous total divides into cells that are not. Decide the cuts before fieldwork. Hunting a significant difference afterward finds one.
Where a behavioral test is available. Pre-orders, deposits, shelf tests and real-money choice experiments answer the commercial question directly, and where one is affordable stated intent is the weaker evidence.
Expectations for recruiting, informing and treating respondents are set by the ESOMAR code and guidelines, the AAPOR standards and ethics materials and the Insights Association code of standards. A favorable intent result is also not substantiation for an advertising claim: the Federal Trade Commission's advertising guidance sets what evidence a claim requires, and consumer enthusiasm is not it.
The organizational limit decides whether any of this matters. A purchase intent number arrives in a room where people have already chosen a favorite. A single percentage is easy to argue with, and easy to weaponize. Agreeing before fieldwork what score triggers what decision converts it from ammunition into evidence.
Teams working out where intent belongs among the other diagnostics will get more from a packaging testing overview and a qualitative research primer than from further refinement of the scale.

