Cluster vs stratified sampling: the difference

By Jude Wallis · Published

Strata are built to be alike inside, and you take a random sample within every stratum. Clusters are built to be a mix inside, each one a small copy of the whole population, and you take a few whole clusters and measure everyone in them. Stratifying buys precision, clustering buys cost.

AP Statistics: Unit 1 (topics 1.11 Random Sampling). Stratified and cluster random sampling are both part of Unit 1 topic 1.11, Random Sampling, in the Fall 2026 AP Statistics course. Unit 1 is 20% to 30% of the multiple-choice section. Questions on this pair usually hand you a scenario, so read what the randomness selected and how many of the groups appear.

Cluster vs stratified sampling: the short answer

Both methods split the population into groups before anyone is selected, and that shared first step is the whole source of the confusion. What separates them is how the groups are built and what the randomness then picks.

A stratified random sample divides the population into strata that are alike inside, then takes a simple random sample within every stratum and combines the pieces into one sample. A cluster random sample divides the population into clusters that are each a mix, ideally a small copy of the whole population, then takes a simple random sample of clusters and measures everyone in the ones drawn.

Two sentences hold the entire distinction, and students reverse both. Strata are homogeneous, clusters are heterogeneous. You sample inside every stratum, and you sample only some of the clusters. Get either sentence backwards and the rest of the topic stops making sense.

The reasons are different too. Stratifying is chosen for precision, because forcing every stratum into the sample takes the differences between strata out of the estimate. Clustering is chosen for cost, because you can run it without a roster of individuals and without traveling to scattered people, and you usually pay for that in extra variability.

Ask two questions about the grouping

Any way of splitting a population can be judged with two questions.

  1. Are the individuals inside one group alike on what you are measuring?
  2. Are the groups alike to one another?

Good strata answer yes to the first and no to the second. Good clusters answer no to the first and yes to the second. The answers are opposite on both counts, which is why one grouping of a population is rarely useful in both roles.

Take a school and the question of how many hours a week students study. Grade level is a natural stratifying variable: if seniors study more like other seniors than like ninth graders, the grades are alike inside and different from each other. Split the same school into homerooms that each mix all four grades and the answers flip. A homeroom holds the school's whole range, and one homeroom looks much like the next. Grades make good strata and poor clusters. Homerooms make good clusters and poor strata.

That gives you the design rule. Group by a trait you already know moves the answer and you have built strata. Build groups that each carry the population's full spread, either by grouping on something unrelated to the answer or by deliberately mixing the trait that moves it, and you have built clusters.

Every stratum, only some clusters

The second reversal is about coverage, and it changes what data you are left holding.

A stratified design reaches into all of its groups. Split a population into 4 strata and sample inside each, and all 4 strata contribute individuals by construction. No draw can leave one out.

A cluster design reaches into a few of its groups. Draw 5 clusters out of 40 and the other 35 contribute nobody. Selection is random, so when the clusters are the same size the method has no systematic lean, but the 35 are simply absent from the data.

That matters when someone asks a follow-up question. After a stratified sample you can report a separate estimate for every stratum, because each one was sampled on purpose. After a cluster sample, any subgroup that lines up with the clusters is either fully in or fully out: if the clusters are 40 schools, then 35 of those schools have no data at all. If a report has to break results out by region or by grade, that is an argument for building strata on the same variable.

The cost case for clustering is worked through in what is a cluster sample, and the choice between the two under a real budget is in how to choose a sampling method.

The differences side by side

FeatureStratified random sampleCluster random sample
Groups are built to beAlike inside, different from each otherMixed inside, alike to each other
The ideal group isA small copy of one subgroupA small copy of the whole population
The randomness selectsIndividuals, inside every groupWhole groups
Groups appearing in the sampleAll of themOnly the ones drawn
Individuals measured in a selected groupA sample of themAll of them
List you need before you drawEvery individual, tagged by stratumThe clusters only
Main payoffA less variable estimateLower cost, no full roster
Usual priceYou need a grouping trait that mattersMore sampling variability at the same size

Read the first row and the third row together. Homogeneity points inward for strata and across groups for clusters, and the randomness works on people in one case and on groups in the other. Every other row follows from those two.

Precision versus convenience

The variability of an estimate comes from what the randomness was allowed to change from one sample to the next, so look at where each design turns the randomness loose.

In a stratified sample the randomness acts inside strata only. Every stratum appears in every possible sample, so the differences between strata never move the estimate. What is left is the variation inside strata, and you built the strata to have little of it. That is why stratifying tends to give a less variable estimate than a simple random sample of the same total size, and why the gain is largest when the strata genuinely differ from one another.

In a cluster sample the randomness acts on the clusters. The differences between clusters become the main thing moving the estimate from sample to sample. You build clusters to resemble each other so that this matters little, but real clusters, meaning city blocks, schools, and homerooms, usually do differ, so a cluster sample of nn people usually carries more variability than a simple random sample of nn people. The convenience is real and so is the price.

The worked example below runs both designs on the same 16 students. The stratified estimate can never miss the true mean by more than 0.5 hours, while the cluster estimate misses it by 2 hours a third of the time. Nothing about the population changed. Only the grouping the randomness was allowed to act on changed.

The classic mix-up and how to avoid it

The mix-up is a straight swap: students say strata should be diverse and clusters should be alike inside. The word stratified probably encourages it, since layers sound like different kinds of things stacked up. Anchor on the reason instead. A stratum exists because you already know a trait that moves the answer, and grouping by such a trait is what makes a group alike inside.

To name a method from a scenario, read it twice.

  1. What did the randomness select, individuals or groups? Individuals point to stratified, groups point to cluster.
  2. How many of the groups show up in the data? All of them means stratified, a few of them means cluster.

Three nearby distinctions are worth keeping straight. Neither method is a simple random sample, because in both of them many groups of size nn can never occur, which the stratified random sample and cluster sample entries work out in full. Strata are not blocks, because blocking groups experimental units before treatments are assigned, a contrast drawn in blocking vs stratifying. And a design that draws whole groups and then samples inside the ones it drew is a multistage sample, not a cluster sample, because a cluster sample measures everyone in a chosen cluster.

One class, both designs, and what each one measures

A statistics elective has 16 students, 4 in each of grades 9 through 12. Their weekly study hours are 3, 4, 4, 5 in grade 9; 5, 6, 6, 7 in grade 10; 7, 8, 8, 9 in grade 11; and 9, 10, 10, 11 in grade 12. Two plans each measure 8 students. Plan S takes a simple random sample of 2 students from each grade. Plan C takes a simple random sample of 2 of the 4 grades and measures all 4 students in each. Name each design and find how far each plan's estimate of the mean can land from the truth.

  1. Find the truth first. The grade totals are 3+4+4+5=163 + 4 + 4 + 5 = 16, 5+6+6+7=245 + 6 + 6 + 7 = 24, 7+8+8+9=327 + 8 + 8 + 9 = 32, and 9+10+10+11=409 + 10 + 10 + 11 = 40, so the class total is 16+24+32+40=11216 + 24 + 32 + 40 = 112 hours over 16 students and the population mean is 11216=7\frac{112}{16} = 7 hours. The four grade means are 4, 6, 8, and 10.

  2. Name the designs. Plan S samples individuals inside every grade, so the grades act as strata and Plan S is a stratified random sample. Plan C picks whole grades at random and measures everyone in them, so the same grades are being used as clusters and Plan C is a cluster random sample.

  3. See what Plan S can measure. Every grade contributes exactly 2 students, so the mean of the 8 is the average of the four grade sample means. The two smallest in each grade give 3+42=3.5\frac{3 + 4}{2} = 3.5, 5+62=5.5\frac{5 + 6}{2} = 5.5, 7+82=7.5\frac{7 + 8}{2} = 7.5, and 9+102=9.5\frac{9 + 10}{2} = 9.5, for an estimate of 3.5+5.5+7.5+9.54=6.5\frac{3.5 + 5.5 + 7.5 + 9.5}{4} = 6.5. The two largest in each grade give 4.5, 6.5, 8.5, and 10.5, for 4.5+6.5+8.5+10.54=7.5\frac{4.5 + 6.5 + 8.5 + 10.5}{4} = 7.5. Every possible Plan S estimate therefore lies between 6.5 and 7.5.

  4. See what Plan C can measure. Each selected grade contributes all 4 of its students, so the mean of the 8 is the average of the two selected grade means. There are (42)=6\binom{4}{2} = 6 equally likely pairs of grades, giving 4+62=5\frac{4 + 6}{2} = 5, 4+82=6\frac{4 + 8}{2} = 6, 4+102=7\frac{4 + 10}{2} = 7, 6+82=7\frac{6 + 8}{2} = 7, 6+102=8\frac{6 + 10}{2} = 8, and 8+102=9\frac{8 + 10}{2} = 9.

  5. Compare the two. Neither design is biased: Plan C's six estimates average to 5+6+7+7+8+96=7\frac{5 + 6 + 7 + 7 + 8 + 9}{6} = 7, and Plan S's possible estimates are symmetric about 7. The spreads are nowhere near each other. Plan S is never off by more than 0.5 hours. Plan C lands away from 7 in 4 of its 6 samples and is off by a full 2 hours in 2 of them, a probability of 26=13\frac{2}{6} = \frac{1}{3}.

  6. See why. The grades were built to be alike inside, which is exactly what makes them good strata and bad clusters. Plan C lets the randomness decide which grades count at all, so the large gaps between grade means turn into the sampling error. Plan S never lets that happen, because all four grades are in every sample.

  7. Fix Plan C by rebuilding the clusters, not the method. Suppose the teacher regroups the 16 into 4 advisories of 4 students, each holding one student from each grade: advisory A has 3, 5, 7, 9 hours, B has 4, 6, 8, 10, C has 4, 6, 8, 10, and D has 5, 7, 9, 11. The advisory means are 6, 7, 7, and 8. Drawing 2 advisories gives 6+72=6.5\frac{6 + 7}{2} = 6.5 twice, 6+82=7\frac{6 + 8}{2} = 7 and 7+72=7\frac{7 + 7}{2} = 7, and 7+82=7.5\frac{7 + 8}{2} = 7.5 twice, so the cluster estimate now lands within 0.5 of the truth every time. Same population, same method, clusters that mirror the class instead of splitting it.

Plan S is a stratified random sample, and its estimate always lands between 6.5 and 7.5 hours. Plan C is a cluster random sample that uses the grades as clusters, and its estimate ranges from 5 to 9 hours, missing the true mean of 7 by 2 hours a third of the time. Rebuilding the clusters as mixed-grade advisories, which is what a cluster is supposed to be, brings the cluster estimate back inside 6.5 to 7.5.

Frequently asked questions

Should clusters be alike inside or alike to each other?

Alike to each other, and mixed inside. The ideal cluster is a small copy of the whole population, which leaves the clusters resembling one another, so any few of them can stand in for the rest. Strata are the reverse: alike inside and different from each other. Clusters that turn out to be alike inside are the case where the design loses the most precision.

Can the same groups be used as strata and as clusters?

Rarely, because the two roles ask for opposite things. Groups built around a trait that moves the answer are alike inside, which makes them good strata and poor clusters. Groups that each carry the population's full range are good clusters and poor strata. A school's grades and its mixed-grade homerooms are those two roles on one population.

Is a cluster sample always less precise than a stratified sample?

No. The penalty depends on how closely the clusters mirror the population. Clusters that really do hold the same mix as the whole can come close to a simple random sample of the same size. Clusters that are alike inside are the bad case, and real clusters such as city blocks or schools usually differ enough from one another to cost some precision.

Why does a stratified sample need a list of individuals when a cluster sample does not?

Because of where the randomness acts. Stratifying draws individuals inside every stratum, so before you can draw you need to know who is in the population and which stratum each person belongs to. Clustering draws whole clusters, so up front you need only a list of the clusters, and you enumerate individuals afterwards, inside the few clusters you selected.