Can you generalize these results? (scope of inference)
By Jude Wallis · Published
Two separate permissions. Random assignment of treatments lets you say the treatment caused the difference. Random selection from a population lets you extend the finding to that population. A study can have both, one, or neither, and your conclusion is capped by whichever one is missing.
AP Statistics: Unit 1 (topics 1.10 The Investigative Question Revisited and Data Collection, 1.13 Experimental Design). In the Fall 2026 course, topic 1.10 asks you to justify the appropriateness of generalizations for a statistical study (skill 2.B) and topic 1.13 asks you to justify conclusions from a well-designed experiment, which together are the scope-of-inference skill tested throughout Unit 1.
Two random acts, two different permissions
Scope of inference is two questions asked one after the other, and the answers come from two different design choices.
Random assignment decides which units receive which treatment. When treatments are randomly assigned to experimental units, the potential for confounding variables is reduced, so a difference in the response between the groups can be attributed to the treatments. That is what buys a cause-and-effect conclusion.
Random selection decides who is in the study at all. When the units in a sample are randomly selected from a population, it is appropriate to make generalizations about the entire population they were drawn from. When they are not, generalizations are appropriate only to a population of individuals similar to those used in the study.
These two choices are independent. You can randomize the assignment without randomizing the selection, and you can randomize the selection without assigning anything at all. The course warns specifically about mixing the two words up: write about random selection when you generalize to a population, and about random assignment when you explain a causal claim.
The 2x2 the exam tests
| How the study was built | Treatments randomly assigned | No random assignment |
|---|---|---|
| Units randomly selected from a population | Cause and effect, generalized to that population | Association only, generalized to that population |
| Units not randomly selected | Cause and effect, but only for units like those in the study | Association only, and only for units like those in the study |
Read it as two independent dials. The columns set what kind of relationship you may claim. The rows set who the claim applies to.
Nothing in one dimension repairs a gap in the other. A convenience sample of 50,000 people still cannot be generalized, and a randomized experiment on 30 volunteers still supports causation for those kinds of volunteers. That is the point people lose most often, and it is why does a bigger sample fix bias answers no.
Both randomizations: the strongest conclusion, and the rarest
A logistics company employs 4,000 warehouse workers. It randomly selects 200 of them from the payroll list, then randomly assigns 100 to a new shift-swapping app and 100 to the existing paper sign-up sheet for one quarter, and records how many extra shifts each worker picks up. The app group picks up more.
Both dials are set. Random assignment means the two groups of 100 were built to be alike on everything except the tool, so the difference in shifts is attributable to the app. Random selection from the payroll list means the 200 stand in for all 4,000 workers, so the conclusion extends to the whole workforce.
This combination is uncommon in practice, and the course says why. It may be unethical or simply difficult to randomly select experimental units to take part in an experiment, so most experiments recruit volunteers instead. When a study does manage both, say both, and cite the matching randomization for each half of the claim.
Random assignment without random selection: the honest case
A researcher recruits 90 volunteers from a university psychology participant pool, randomly assigns 45 to a 9 p.m. screen curfew and 45 to no curfew for two weeks, and measures nightly sleep from a wearable. The curfew group sleeps longer.
The causal conclusion survives intact. Random assignment did its job: the two groups were built to be alike on everything except the curfew, so the difference in sleep is attributable to the curfew. What does not survive is the audience.
The 90 volunteers were not randomly selected from any defined population, so the finding applies to a population of individuals similar to those in the study, meaning undergraduates who signed up for a sleep experiment. It does not apply to adults in general, to shift workers, or to teenagers.
This is the most common shape in real research, and it is not a flaw to apologize for. Most experiments run on people who agreed to be in them, because drafting people into a treatment at random is usually impossible and often unethical. State the causal claim, state the narrow audience, and stop. Designing such a study is covered in how to design an experiment.
Random selection without random assignment: generalize, but not to a cause
A state agriculture office takes a simple random sample of 250 of the 4,000 licensed beekeepers in the state and asks about winter hive losses and hive insulation. Beekeepers who insulate report lower losses.
Here the audience is wide and the claim is weak. Because the 250 were randomly selected from the full list, generalizing to all 4,000 licensed beekeepers is appropriate. Because nobody assigned insulation, the study cannot say insulation caused the lower losses.
The reason is confounding. Beekeepers chose for themselves whether to insulate, so a variable associated with both insulating and hive survival, such as years of experience or how aggressively someone treats for mites, offers a rival explanation that the study cannot rule out. A confounding variable has to be associated with both the explanatory and the response variable, and in observational studies it usually is.
Every well-run survey lands in this cell, which is why survey results are reported as associations. The distinction between assigning conditions and merely recording them is drawn in experiments vs observational studies, and the reasoning trap it creates is in correlation vs causation.
Neither: describe the sample and stop
A gym posts a sign-up sheet for a new stretching class and later compares the flexibility of the 40 members who signed up with 40 regulars who did not. The stretchers are more flexible.
Both permissions are gone at once. There is no random selection, because participants put themselves on the list, and there is no random assignment, because they also chose their own group. You can describe what was observed in these 80 members, and that is the whole of what the study supports.
The rival explanations write themselves. People who join a stretching class may already be more flexible, younger, or more motivated than people who do not, so the observed gap has as many candidate causes as there are differences between the two groups. Self-selection into groups is the mechanism that produces confounding, which is why nonrandom sampling methods carry potential bias by construction. Picking a defensible method instead is the subject of how to choose a sampling method.
How to write the scope-of-inference answer
A scope-of-inference answer is two sentences, each with its reason attached. Both must name the actual study and the actual population, not the general rule.
- The causation sentence. Either: because the treatments were randomly assigned, it is reasonable to conclude that the treatment caused the difference in the response. Or: because no treatments were assigned, the study shows an association but cannot establish that one variable caused the other.
- The generalization sentence. Either: because the units were randomly selected from the population, the results can be generalized to that population. Or: because the units volunteered rather than being randomly selected, the results extend only to individuals similar to those in the study.
Two habits protect the points. Give a clear yes or no first and then the reason, because a hedge with no decision earns nothing. Then cite the right randomization for the right half: the causal sentence cites random assignment, and the generalization sentence cites random selection.
Naming a placebo, blinding, a control group, or a large sample as the reason for a causal claim does not work, because none of those balance the extraneous variables across groups. Only random assignment does that. Scope of inference sits in topic 1.10 and again in topic 1.13, and you can rehearse the wording in experimental design practice and sampling methods practice. The official framework is at AP Central.
Worked answer: a randomized experiment on volunteers
A researcher recruits 90 volunteers from a university psychology participant pool. She randomly assigns 45 to a 9 p.m. screen curfew and 45 to no curfew for two weeks, and records nightly sleep minutes from a wearable. The curfew group averages 27 more minutes of sleep per night, a statistically significant difference. Write the scope-of-inference answer.
Check for random assignment. The researcher used a random mechanism to split the volunteers, and the counts confirm everyone was used: volunteers. Treatments were randomly assigned.
Check for random selection. The 90 came from a pool of people who volunteered, not from a random draw out of a defined population. The sample was not randomly selected.
Set the causal verdict from the assignment answer. Random assignment reduces the potential for confounding by spreading extraneous variables about equally across the two groups, so the 27-minute difference can be attributed to the curfew rather than to some other difference between the groups.
Set the audience from the selection answer. Because the participants volunteered, the finding applies to a population of individuals similar to them, university students who signed up for a sleep experiment, and not to adults in general.
Write both sentences using the study's own nouns. Say curfew and sleep minutes and psychology pool, not treatment and response and population.
Yes, it is reasonable to conclude that the 9 p.m. screen curfew caused the increase in sleep, because the researcher randomly assigned the 90 volunteers to the two conditions, which makes the two groups similar on other variables that affect sleep. No, the result should not be generalized to all adults, because the 90 participants volunteered from a university psychology pool rather than being randomly selected from a population; the conclusion extends only to students similar to those who took part.
Worked answer: a random sample with no assigned treatment
A state agriculture office takes a simple random sample of 250 of the 4,000 licensed beekeepers in the state and asks each one about winter hive losses and whether they insulate their hives. Among beekeepers who insulate, 18% of hives were lost over the winter; among those who do not, 29% were lost. Write the scope-of-inference answer.
Check for random assignment. Nobody assigned insulation; each beekeeper decided independently. This is an observational study, so there is no random assignment.
Check for random selection. The 250 were drawn as a simple random sample from the list of all 4,000 licensed beekeepers, a fraction of , or 6.25% of the population. Selection was random.
Size the finding before judging it. The gap is percentage points, which is a real difference worth explaining, not a rounding artifact.
Set the causal verdict. With no assigned treatment, that 11 point gap is an association. A confounding variable, one associated with both insulating and hive loss, could explain it: beekeepers who insulate may also be more experienced or more thorough about treating for mites.
Set the audience. The sample was drawn at random from the full list of licensed beekeepers, so generalizing to all 4,000 of them is appropriate. It does not extend to unlicensed or hobbyist beekeepers who are not on that list.
Notice the shape. This is the mirror image of the volunteer experiment: a wide audience with a weak claim, where the experiment had a narrow audience with a strong claim.
No, the office cannot conclude that insulating hives causes lower winter losses. Beekeepers chose for themselves whether to insulate, so the 11 percentage point difference is an association, and a confounding variable such as overall management experience is plausibly linked to both insulating and hive survival. Yes, the results can be generalized to all 4,000 licensed beekeepers in the state, because the 250 surveyed were a simple random sample from that population.
Frequently asked questions
If the subjects were volunteers, is the experiment worthless?
No. Random assignment still balances extraneous variables across the treatment groups, so the causal conclusion holds for the units in the study. What shrinks is the audience: the result applies to individuals similar to those who volunteered. Most published experiments are in exactly this position.
Can a very large sample make up for not being randomly selected?
No. Size controls variability, not bias. A nonrandom sampling method has a systematic error built into who gets in, so a larger sample estimates the wrong quantity more precisely. Only random selection from the population you care about buys generalization.
What counts as "individuals similar to those in the study"?
Describe the group that was actually recruited, using the details the study gives you: undergraduates in one psychology pool, patients at one clinic, plots in one greenhouse. Vague phrases like similar people earn little. Naming the recruitment channel is usually the clearest way to bound the claim.
Which randomization do I cite for which claim?
Random assignment for the causal claim, random selection for the generalization claim, and never the other way round. A response that cites a large sample, a placebo, or blinding as the reason a treatment caused an effect misses the point, because none of those balance extraneous variables across the groups.
Can an observational study ever establish cause and effect?
Not on its own. Without assigned treatments, any variable associated with both the explanatory and the response variable offers a rival explanation, and the study has no way to rule it out. Observational studies can point at a relationship worth testing, and a randomized experiment is what settles it.