Home Feeds Careers Get in Touch

Cluster Sample vs Stratified Sample and When to Use Each 2026

cluster sample vs stratified sample cluster sampling vs stratified cluster sampling vs stratified sampling cluster versus stratified sampling stratified sampling vs cluster sampling stratified sampling cluster sampling multistage cluster sampling design effect intraclass correlation
Cluster Sample vs Stratified Sample and When to Use Each 2026

TL;DR

  • Stratified sampling splits a population into groups and samples inside every one, which usually shaves a little variance.
  • Cluster sampling samples whole groups and ignores the rest, which usually adds a lot.
  • The design effect measures the gap: at 2.0, you field 800 interviews for the precision of 400.
  • Stratify when you can list the population and groups differ. Cluster when you cannot list it and fieldwork has to travel.

Last updated: 17 September 2026

Quick Answer: Stratified sampling buys precision, cluster sampling buys cost, and the design effect measures the trade. The deciding difference is where the similarity sits: stratify when groups differ and you can list the population, cluster when you cannot and fieldwork must travel. The United Nations default design effect of 2.0 means fielding 800 interviews for the precision of 400.

Stratified sampling divides a population into groups and samples inside every group. Cluster sampling divides it into groups, samples only some of them, and never visits the rest.

One published case makes it concrete. A trial powered at 80% with 128 patients asked four physicians to recruit 32 each, then came back rejected for inadequate power at 61%: 32 patients per physician is a cluster, and nobody had adjusted for it.

Why Does One Design Cut Error and the Other Cut Cost?

Stratification reduces variance when strata are internally homogeneous and differ from each other. Clustering does the opposite: people in the same cluster resemble each other, so each extra interview inside one tells you less than a fresh one does.

The logic goes back to Cochran, quoted in Statistics Canada's own methods manual. If each stratum is homogeneous, a precise estimate of its mean comes from a small sample in it, and those estimates combine into a precise estimate for the population.

The design effect is the whole comparison in one number.

  • Design effect = 1 + (m − 1) × ρ, where m is interviews per cluster and ρ is the intraclass correlation.
  • Effective sample size = total interviews ÷ design effect.

Run the trial through it. At 32 patients per physician and an intraclass correlation of 0.017, the design effect is 1.527, so 128 enrolled patients carried the information of 84. In budget terms, a design effect of 1.5 means fielding 600 interviews for the precision of 400; at 2.0, 800.

The United Nations handbook on household survey samples sets that default, with 1.5 for designs following its cluster rules. It is also blunt about the asymmetry: stratification decreases sampling variance only to a small degree, while clustering increases it considerably.

What Is Stratified Sampling and What Does an Example Look Like?

Stratified sampling divides the frame into mutually exclusive strata, then draws an independent sample inside every one.

A worked example. A beverage brand wants a usage and attitude study, n = 1,200, off a customer file of 240,000 buyers carrying purchase frequency: heavy buyers 10% of the file, medium 30%, light 60%.

  • Proportionate allocation gives 120 heavy, 360 medium, 720 light. National numbers are clean, but the heavy-buyer cell, the one the category manager wants, carries a margin of error near ±8.9 points on a 50% reading.
  • Disproportionate allocation gives 400 in each cell. Heavy-buyer precision improves to roughly ±4.9 points.

Restoring the national estimate needs weights of 0.3 for heavy, 0.9 for medium and 1.8 for light. Unequal weights carry a design effect of their own, 1.38 here, so 1,200 interviews behave like about 870 nationally. Kish's formula treats the overall design effect as the product of the unequal-weighting effect and the clustering effect, so a disproportionately stratified cluster sample pays both.

How Does a Cluster Sample Work, and What Does It Cost?

Cluster sampling draws whole naturally occurring groups, then measures everyone inside the selected ones (one-stage) or a sample of them (two-stage).

What does a cluster sampling example look like? A retailer wants 1,000 shopper intercepts across a national grocery footprint. No list of grocery shoppers exists. A list of stores does. Draw 40 stores with probability proportional to footfall, then intercept 25 shoppers in each. At ρ = 0.03 and 25 per store, the design effect is 1.72, so 1,000 intercepts read like 581.

Now change the shape. Eighty stores at 13 intercepts is 1,040 interviews, a design effect of 1.36, an effective sample near 765. Four percent more interviewing, 184 more effective interviews, twice the recruitment and travel. The UN handbook points the same way: for 12,000 households, 600 clusters of 20 beats 400 of 30.

Where Do Multistage Designs Fit?

Multistage sampling means selection happens in more than one step, and multistage cluster sampling means those steps are clusters: regions, then blocks, then households. Stratified says how the frame is divided at a step, so almost every national survey is both, and every stage adds its own intraclass correlation.

Which Design Should You Use for a Consumer Study?

Rows are sorted alphabetically, so the order carries no ranking.

Design What it buys What it costs When it is the right call
Cluster sampling (one-stage) Fewer fieldwork locations Variance inflation from clustering In-person work where travel dominates
Multistage cluster Reach without an individual list A correlation from every stage National studies, no address register
Non-probability panel with quotas Speed, census-matched profile No selection probabilities, no design-based variance Concept screens, directional reads
Stratified sampling Total precision plus subgroup bases Effective sample, if allocation is disproportionate Customer files, B2B universes

What Do These Designs Assume That Commercial Fieldwork Rarely Delivers?

Both designs assume a sampling frame that enumerates the population. AAPOR's description of probability sampling is exact: a frame covering all or almost all of the population of interest, with individuals selected from it at random. Consumer research rarely has one.

What a commercial brief calls a stratified sample is usually interlocking quotas on an opt-in panel, a different object with a different error structure. Where that panel's people come from decides its error, as where good respondents come from sets out.

Pew Research Center asked opt-in respondents whether they held a nuclear submarine license. Twelve percent of those under 30 said yes, as did 24% of opt-in cases claiming to be Hispanic, against a true share that rounds to zero. The distortion concentrates in exactly the cells a quota grid exists to protect.

Who Can the Channel Actually Reach Inside a Stratum?

Even with a real frame, the channel decides who inside a stratum or a cluster can be interviewed. The ITU's Facts and Figures 2025 reports almost three-quarters of the world online while urban and rural divides narrow without closing.

A stratum of rural households in a middle income band is not sampled by a browser link. Its connected, private-room, spare-device subset is, and it differs on the variables the stratification existed to control.

Alchemic designs the sample and fields it, recruiting through managed fieldwork or bring your own across 14 markets including the USA and the UK. Respondents pick the channel: a web link, an interview running natively inside WhatsApp with no link and no app, or an outbound AI phone call.

Where Neither Design Saves a Study

Neither design touches nonresponse. Both govern who gets selected; who agrees is a separate error term, and AAPOR's standard definitions exist because the profession needed one vocabulary for it. A perfectly stratified sample with 4% response is a convenience sample in a costume.

Cluster designs cannot outrun a bad frame: a three-year-stale outlet list hands the sample every closure it missed.

For a 30-interview qualitative study the design effect is not the binding question; whether those 30 span the segments you intend to describe is. Sample validity in AI-moderated research covers it, as does reaching respondents without smartphones.

For a population with no phone, no internet and no list, where measurement must happen in the home, an in-person area-probability design run by a field house with local enumerators still beats every remote alternative, including anything Alchemic fields. The overview of types of market research maps where these designs sit.

Frequently Asked Questions

What is a good design effect to assume before fieldwork starts?
Use 2.0 unless prior data from a similar survey says otherwise. The United Nations handbook on household survey samples sets that default for sample size calculations, and allows a figure nearer 1.5 for designs using many small clusters. A design effect can only be measured afterwards, from the data.
Is stratified sampling always more accurate than simple random sampling?
No. Stratification improves precision only for variables correlated with the stratifying variable. Stratify a business survey by employee count and sales estimates get sharper, while estimates of employee age do not, because headcount predicts one and not the other. For uncorrelated measures, a stratified design can be slightly less efficient.
How many clusters does a cluster sample need?
More than most plans assume. The design effect grows with interviews per cluster, not with cluster count, so adding clusters raises effective sample size faster than adding interviews inside existing ones. In the published trial, four physicians would have needed 320 patients for 80% power, while sixteen needed 160.
What is the difference between stratified sampling and quota sampling?
Stratified sampling selects randomly inside strata drawn from a frame, so every unit has a known selection probability and sampling error can be calculated. Quota sampling fills demographic cells with whoever is available and willing. The two profiles can look identical while the inferential standing is not.
Do stratified and cluster samples need to be weighted?
Usually yes. Proportionate stratification can be self-weighting, but disproportionate allocation, unequal cluster selection probabilities and differential response all require weights. Weights themselves carry a design effect from their variation, which multiplies with the clustering effect rather than replacing it, so weighting is part of the precision budget.

About the Author

Sreenadh Narayanan is the founder of Alchemic, an AI-powered consumer research platform used for ad testing, concept testing and brand tracking. He writes Alchemic's guides on qualitative research and research methods, covering interview design, sample sizes and how teams turn customer conversations into decisions.