Experiment vs observational study: what to conclude
By Jude Wallis · Published
An experiment imposes treatments on units, and when it also uses random assignment it can support a cause-and-effect conclusion. An observational study only records variables, so it shows association, not cause. Random selection is a separate step that lets you generalize to the population.
AP Statistics: Unit 1 (topics 1.10 The Investigative Question Revisited and Data Collection, 1.11 Random Sampling, 1.12 Potential Problems with Sampling, 1.13 Experimental Design). In the Fall 2026 course, Unit 1 topics 1.10 through 1.13 cover data collection, where random assignment in an experiment justifies a cause-and-effect conclusion and random selection justifies generalizing to the population.
Experiment vs observational study: the core difference
The difference between an experiment and an observational study comes down to one question: did the researcher impose a treatment? In an experiment, the researcher assigns conditions, called treatments, to experimental units to answer a question about a population. In an observational study, no treatment is imposed, and the researcher only records the values of the variables of interest.
A few terms make the comparison precise. The explanatory variable (also called a factor in an experiment) is the variable whose levels you impose in an experiment or simply record in an observational study. In an experiment those imposed levels are the treatments, and an observational study has no treatments. The response variable is the outcome you measure on each unit after any treatment. An experimental unit is the person or object a treatment is assigned to, and when the units are people they are often called subjects or participants.
Here is the quick test. If the sentence describing the study has a researcher deciding who gets what (assigning a drug or a study method), it is an experiment. If the researcher steps back and records what people already do or already are, it is an observational study. A survey is an observational study that collects data from people using a standard set of questions.
Confounding: why an observational study cannot prove cause
An observational study can reveal a strong relationship between two variables, but it cannot rule out a confounding variable. In an observational study, a confounding variable is one that is associated with both the explanatory variable and the response variable, so it offers an alternative explanation for the relationship you see.
Suppose records show that adults who drink more coffee sleep worse. Coffee intake is the explanatory variable and sleep quality is the response. Stress could be a confounder, because stressed people may drink more coffee and also sleep worse on their own. Since you never controlled who drank coffee, you cannot separate the effect of coffee from the effect of stress.
An experiment attacks this problem directly. In a well-designed experiment the potential for confounding is reduced, because the design controls who receives each treatment. That control is exactly what an observational study gives up, which is why an observational study supports association but not cause.
Random assignment vs random selection: two kinds of randomness
AP Statistics uses randomness in two separate places, and keeping them apart is the whole point.
Random assignment happens inside an experiment. Once you have your experimental units, you use chance (a random number generator or drawing numbered slips) to decide which unit gets which treatment. Its purpose is to create treatment groups that are as similar as possible with respect to extraneous variables, the outside factors that could affect the response, so those factors end up balanced across the groups on average. That balance is why random assignment supports a cause-and-effect conclusion.
Random selection (random sampling) happens when you choose who is in the study in the first place. You draw units from the population using a random mechanism, such as a simple random sample. Its purpose is a representative sample, so random selection is what lets you generalize your results to the population.
They answer different questions. Random assignment earns you causation, and random selection earns you generalization. A study can have one, both, or neither.
Scope of inference: what you can actually conclude
Because the two kinds of randomness are independent, four combinations are possible. This scope-of-inference table shows what each one lets you claim. Read the rows for how the units were chosen and the columns for whether treatments were randomly assigned.
| How units were chosen | Random assignment (experiment) | No random assignment (observational) |
|---|---|---|
| Random sample from the population | Cause and effect, generalized to the population | Association only, generalized to the population |
| Volunteers or convenience sample | Cause and effect, but only for units like those studied | Association only, and only for units like those studied |
Two rules generate the whole table. Random assignment of treatments allows a cause-and-effect conclusion. Random selection of units allows generalization to the population they were drawn from, and without it you can generalize only to individuals similar to those in the study.
Most real experiments live in the bottom-left cell: treatments are randomly assigned, so cause is fair, but the subjects are volunteers rather than a random sample, so the conclusion applies only to people like them. That is a valid and common place to be, and the AP exam accepts it as long as you state the limit.
How the AP exam tests experiments vs observational studies
This material is Unit 1, topics 1.10 through 1.13, part of the 20% to 30% of the multiple-choice section that Unit 1 carries. It also feeds Statistical Practice 2, Collect Data, which is 20% to 30% of the multiple-choice questions, and it appears in the free-response section, where Question 1 is a multi-focus item on formulating questions and collecting data.
Expect three moves on the exam:
- Classify. Is the scenario an experiment or an observational study? Look for whether a treatment was imposed.
- Critique. Name a specific confounding variable, or point to a missing element of a well-designed experiment: a comparison group, random assignment, replication, or direct control.
- State the scope. Use the two kinds of randomness to say whether cause and generalization are justified.
A common trap is to reward a large sample size with a causal conclusion. Sample size does not create causation, only random assignment does. Another trap is a randomized experiment run on volunteers, where causation is fine but generalization is limited to people like the volunteers.
For the correlation side of the same idea see correlation vs causation, and for the full unit see Unit 1: Exploring One-Variable Data. The official framework is at AP Central.
Classify and critique: coffee and exam scores
A researcher records the daily coffee intake and final exam scores of 500 college students who were randomly selected from one university. Students who drink more coffee tend to score lower. Classify the study, name a confounding variable, and state the scope of inference.
Classify the design. The researcher only recorded coffee intake and scores and imposed no treatment on the students, so this is an observational study (a survey-style record of existing behavior).
Look for a confounder. A confounding variable must be associated with both the explanatory variable (coffee intake) and the response (exam score). Sleep fits: students who sleep less may drink more coffee and also score lower, which offers an alternative explanation for the pattern.
Judge causation. Because coffee was not randomly assigned, confounders like sleep are not balanced out, so the study cannot support the claim that coffee lowers scores. It shows association only.
Judge generalization. The 500 students were randomly selected from the university, so it is appropriate to generalize the association to the population of that university's students, but not beyond it.
Observational study. It shows an association, not causation, because sleep is a plausible confounder, and the result generalizes to that university's student population since the sample was randomly selected.
Classify and critique: a randomized fertilizer trial
A gardening company recruits 40 volunteer gardeners. Each gardener's plot is randomly assigned to receive either a new fertilizer or the standard fertilizer, and the tomato yield is measured. The new-fertilizer plots yield more on average. Classify the study and state what can be concluded.
Classify the design. The researcher imposed a treatment (which fertilizer each plot received) and used chance to assign it, so this is an experiment. Because treatments were assigned completely at random, it is a completely randomized design.
Check the elements. It compares two treatment groups, new versus standard (the standard acts as the control), it uses random assignment, and it has replication because many plots receive each treatment. Those are marks of a well-designed experiment.
Judge causation. Random assignment balances extraneous variables such as soil quality and sunlight across the two groups, so a difference in yield can be attributed to the fertilizer. A cause-and-effect conclusion is justified.
Judge generalization. The 40 gardeners were volunteers, not a random sample from any larger population, so the causal conclusion applies only to gardeners and plots similar to those studied, not to all gardeners.
Experiment (completely randomized design). Random assignment supports the cause-and-effect conclusion that the new fertilizer raises yield, but because the gardeners volunteered, that conclusion generalizes only to units similar to those studied.
Classify and critique: a nationwide randomized study
A researcher takes a simple random sample of 2,000 adults from a national registry, then randomly assigns each person to a 10-minute daily walking program or no program for eight weeks, and measures resting heart rate. The walking group ends with a lower average resting heart rate. What is the scope of inference?
Classify the design. A treatment (walk or no walk) was imposed and randomly assigned, so this is an experiment.
Judge causation. Random assignment is present, so cause and effect is on the table: the lower heart rate can be attributed to the walking program rather than to confounders, which random assignment balances across the groups.
Judge generalization. Random selection is also present, because the 2,000 adults were a simple random sample from the national registry, so the results generalize to that population.
Combine the two. With both randomizations present, this study sits in the strongest cell of the scope-of-inference table, where both cause and generalization are justified.
Both randomizations are present, so you can conclude that the walking program causes a lower resting heart rate and generalize that conclusion to the adult population in the national registry.
Frequently asked questions
Can an observational study ever show cause and effect?
No. Without random assignment, a confounding variable can always offer an alternative explanation for the relationship. An observational study can flag a strong association and motivate an experiment, but the causal conclusion has to come from an experiment that randomly assigns treatments.
What is the difference between random assignment and random selection?
Random assignment decides which treatment each unit gets inside an experiment, and it is what supports a cause-and-effect conclusion. Random selection decides which units enter the study in the first place, and it is what lets you generalize to the population. A study can use one, both, or neither.
Does a larger sample make an observational study able to prove cause?
No. A larger sample gives a more precise estimate, but it does not remove confounding. Ten thousand observations with no random assignment still leave lurking variables in play, so the study shows association, not cause.
Is a survey an experiment or an observational study?
A survey is an observational study. It collects data from people using a standard set of questions but imposes no treatment, so it can describe and relate variables while it cannot establish cause on its own.