How to design an experiment and describe it
By Jude Wallis · Published
Build it on three principles: compare at least two treatment groups, assign treatments to units at random, and put more than one unit in each treatment. Then describe it in four moves: name the treatments, give the exact random mechanism, name what you measure, and say what you compare.
AP Statistics: Unit 1 (topics 1.13 Experimental Design, 1.10 The Investigative Question Revisited and Data Collection). In the Fall 2026 course, topic 1.13 asks you to identify the elements of a well-designed experiment and name the design (skill 2.A) and to justify why a particular design fits a study (skill 2.B), while the vocabulary of experimental units, factors, and treatments is set in topic 1.10.
The three principles a design has to satisfy
A well-designed experiment rests on three things, and a design missing any of them cannot answer the question it was built for.
Comparison. You need at least two treatment groups, and one of them may be a control group. A control group is a set of experimental units created for comparison, and it may receive a different treatment, such as an inactive substance called a placebo, rather than nothing at all. One group measured on its own gives you a number with nothing to hold it against.
Random assignment. Treatments are assigned to experimental units by a random mechanism. Its purpose is to make the treatment groups as similar as possible with respect to extraneous sources of variation, so that each extraneous variable ends up spread about equally across the groups. That balancing is why an experiment can support a cause-and-effect conclusion at all.
Replication. More than one experimental unit is assigned to each treatment. With a single unit per treatment you cannot separate a treatment effect from the ordinary variation between units. Bigger groups make the comparison more precise.
The course lists a fourth element beside these three: direct control of other sources of variation, meaning you keep the settings you are not studying the same from unit to unit. Same room, same duration, same measuring instrument.
Vocabulary the description has to use correctly
A handful of words carry specific meanings, and swapping them costs points.
- An experimental unit is the thing a treatment is assigned to. When the units are people they are called subjects or participants.
- The explanatory variable, or factor, is the variable whose levels you impose. Those levels are the treatments. With two explanatory variables, the treatments are the combinations of their levels.
- The response variable is the outcome measured on each unit after the treatment has been administered.
- An extraneous variable is one believed to affect the response that you are not studying. A confounding variable is tied to the explanatory variable closely enough that you cannot tell which of the two moved the response.
A study counts as an experiment only when the researcher assigns the conditions. If the researcher records what units already do, it is an observational study, and no amount of careful measurement converts it into an experiment. That line is drawn in experiments vs observational studies, and the vocabulary comes from topic 1.10.
The four moves of a full-credit description
Descriptions are scored piece by piece, so one that nails three of these and drops the fourth is not close, it is incomplete. Write all four every time.
- Name the treatments. Say exactly what each group receives, including the control or placebo. A new study app and the current worksheet packet are treatments; the treatment group is not.
- Say how the random assignment happens. Give a mechanism a stranger could carry out: number the units 1 to , where is the total number of experimental units, generate random integers in that range, ignore repeats, and state which numbers go where and where the leftover units land. The word randomly, on its own, describes nothing.
- State what is measured. Name the response variable precisely, with units where they exist: time in seconds, score out of 40, yield in kilograms per plot. Performance is not a response variable.
- Say what will be compared. Finish with the sentence that compares the groups on that response, usually the mean response or the proportion with some outcome.
The fourth move is the one people leave out. They build the randomization carefully, then stop, and the design never says what question it answers. A longer walkthrough of the same write-up for one specific design is in how to describe a completely randomized design.
Write the random assignment as a procedure
Two mechanisms cover everything the exam asks for.
- Random number generator. Number the units 1 to . Generate random integers from 1 to , ignoring repeats, until you have as many distinct numbers as the first group needs. Those units take treatment 1. Repeat for treatment 2, and the units never drawn take the last treatment.
- Physical randomization. Write each unit's number on an identical slip of paper, put the slips in a container, mix thoroughly, and draw the number of slips each group needs in turn.
Two clauses get dropped, and both matter. Say that repeats are ignored, because a generator will hand you the same number twice and the procedure has to know what to do. Say where the leftover units go, or the last group never fills.
Be careful with flipping a coin for each subject. It is genuine random assignment, but it does not deliver the group sizes you just promised, so use the generator or the container whenever you have committed to fixed sizes.
Completely randomized, randomized block, or matched pairs
The three designs differ in what happens before the randomization.
| Design | What happens first | How treatments are assigned | When to use it |
|---|---|---|---|
| Completely randomized | Nothing; all units sit in one pool | Treatments assigned to the units completely at random | Nothing is known about an extraneous variable worth separating out |
| Randomized block | Units are grouped into blocks that are alike on a blocking variable | Treatments assigned at random within each block, so every treatment appears in every block | A known extraneous variable affects the response and you want it out of the comparison |
| Matched pairs | Units are paired on one or more extraneous variables, or each unit serves as its own pair | One treatment assigned at random to each member of the pair, or the order of the two treatments randomized within a unit | There are exactly two treatments and pairing removes most of the unit-to-unit variation |
A blocking variable is a source of extraneous variation in the response. Blocking separates the variation that variable causes from the rest, so treatments can be compared inside a block without that variable getting in the way. A matched pairs design is simply a randomized block design with two treatments and blocks of size two.
Do not add blocks nobody asked for. If the prompt hands you a variable expected to affect the response and asks you to account for it, block on it; if it does not, describe the completely randomized design and stop. Blocking in an experiment and stratifying in a survey are the same idea aimed at different problems, compared in blocking vs stratifying.
Blinding, placebos, and direct control
Blinding keeps knowledge of the treatment from changing either the response or the measurement.
- Single-blind, also called single-masked: the participants do not know which treatment they are receiving, but the researchers who interact with them do, or the other way around.
- Double-blind, also called double-masked: neither the participants nor the researchers who interact with them know which treatment each participant is receiving.
A placebo is an inactive substance given so that the control group has the same experience as the treatment group apart from the active ingredient. The placebo effect is the difference between the average response to a placebo and the average response to no treatment, which is why a placebo group is a better comparison than an untreated group.
Direct control is the easiest element to write and the one most often skipped. Name two or three conditions you hold constant across all units: the same room, the same time of day, the same measuring instrument. Anything held constant cannot vary across the groups, so it cannot explain a difference between them.
Mistakes that cost points
These come up again and again in descriptions that look finished.
- Writing randomly assign with no mechanism attached. This is the most common way to lose the randomization credit.
- Leaving out the group sizes, so a reader cannot tell whether every unit was used.
- Naming a vague response such as improvement instead of score on the same 40-question test.
- Never writing the comparison sentence, so the design gathers data but answers nothing.
- Sorting units before randomizing, for example putting the strongest half in one group. That is not random assignment, and it reintroduces the confounding the design exists to prevent.
- Blocking when the question asked for a completely randomized design, or the reverse.
- Claiming the result generalizes to everyone because the assignment was random. Random assignment supports a causal claim; random selection is what supports generalization, a split worked through in can you generalize these results.
Design questions live in topic 1.13 and are tested in multiple choice and on the free-response question about collecting data. To rehearse the write-up under time pressure, use experimental design practice and the AP Statistics FRQ guide. The official framework is at AP Central.
A full design description you can copy the shape of
A community college wants to know whether a 10-minute guided breathing exercise before an exam lowers test anxiety. Seventy-two students in an introductory course agree to take part. All of them will sit the same 90-minute midterm and then complete the same 20-item anxiety questionnaire, scored from 0 to 60, with higher scores meaning more anxiety. Describe a completely randomized design.
Identify the units and the treatments. The experimental units are the 72 students. Treatment 1 is a 10-minute guided breathing recording before the exam; treatment 2, the control, is sitting quietly in the same room for 10 minutes with no recording.
Set the group sizes. Two treatments split evenly gives students per group.
Write the random assignment as a procedure. Number the students 1 to 72, use a random number generator to produce random integers from 1 to 72, ignore any repeat, and continue until 36 distinct numbers appear. Those 36 students get the breathing recording, and the remaining students sit quietly.
Add direct control. Both groups use the same room, start at the same time, sit the same 90-minute midterm, and complete the questionnaire immediately after handing the exam in.
Add blinding where it is possible. The students necessarily know which activity they did, but the staff member who scores the questionnaires is not told which group each student was in, which makes the scoring single-blind.
Name the response. Record each student's score on the 20-item questionnaire, from 0 to 60.
State the comparison. Compare the mean questionnaire score of the 36 students who did the breathing exercise with the mean score of the 36 students who sat quietly.
Audit the three principles before moving on: comparison (two treatment groups, one of them a control), random assignment (the generator, with repeats ignored and the leftover units placed), and replication (36 units per treatment, far above one).
Number the 72 students 1 to 72. Use a random number generator to generate random integers from 1 to 72, ignoring repeats, until 36 distinct numbers appear; those 36 students listen to the 10-minute guided breathing recording before the midterm, and the other 36 sit quietly for 10 minutes in the same room. All 72 sit the same 90-minute midterm at the same time in the same room and complete the same 20-item anxiety questionnaire immediately afterward, scored from 0 to 60 by a staff member who is not told which group each student was in. Record each student's score, then compare the mean anxiety score of the 36 students who did the breathing exercise with the mean score of the 36 who sat quietly.
The same study as a randomized block design
Before the study starts, the instructor notices that students who already report high everyday anxiety may respond very differently from students who do not. On a baseline anxiety survey, 36 of the 72 students score above the class median and 36 score at or below it. Redesign the study as a randomized block design, and say what the blocking buys.
Choose the blocking variable. Baseline anxiety is an extraneous source of variation: it is expected to affect the questionnaire score, but it is not the variable being studied.
Form the blocks. Block A is the 36 students above the median, block B is the 36 at or below it, and students within a block are alike on baseline anxiety.
Randomize inside each block, never across blocks. Number the students in block A from 1 to 36, use a random number generator to draw 18 distinct integers from 1 to 36 while ignoring repeats, and give those 18 the breathing recording; the other students in block A sit quietly. Run the identical procedure separately inside block B.
Check the counts. Each block splits into per treatment, so each treatment is used times overall and students are assigned, matching the total.
Measure the same response under the same conditions. Record every student's score on the same 0 to 60 questionnaire, with the same room, timing, and blinded scoring as before.
Compare within blocks first. Compare the mean score of the breathing group against the quiet group inside block A, do the same inside block B, then combine the two block comparisons.
State what the blocking buys. Differences in baseline anxiety no longer sit inside the treatment comparison, because every comparison is made among students who started at a similar level, so a real effect of the breathing exercise is easier to detect.
Split the 72 students into a high-baseline block of 36 and a low-baseline block of 36. Inside each block, number the students 1 to 36 and use a random number generator, ignoring repeats, to pick 18 students for the breathing recording; the other 18 in that block sit quietly. Measure the same 0 to 60 anxiety score under identical conditions, compare the two treatment means inside each block, then pool the block comparisons. Blocking pulls baseline anxiety out of the treatment comparison, so this design detects a smaller true effect than the completely randomized version would.
Frequently asked questions
Is writing "randomly assign the subjects to the treatments" enough?
No. That names the idea without describing it, and it is the most common place descriptions lose credit. Give a procedure a stranger could carry out: number the units, generate random integers in that range, ignore repeats, say which numbers go to which treatment, and say where the units never drawn end up.
Does every experiment need a control group?
It needs at least two treatments to compare, and one of them may be a control group receiving a placebo or the current standard. Comparing two active treatments also satisfies the comparison principle. What you cannot do is run one group and call it an experiment, because there is nothing to compare it against.
How many units per treatment counts as replication?
Replication within an experiment means more than one experimental unit is assigned to each treatment, so two per treatment technically satisfies the definition. In practice you want considerably more, because larger groups make the comparison between treatments more precise and make a real effect easier to detect.
What is the difference between blocking and stratifying?
Blocking happens in an experiment: you group experimental units by a variable that affects the response, then randomly assign treatments within each block. Stratifying happens in a survey: you group the population by a shared trait, then take a random sample within each stratum. Same idea, different job, and the words are not interchangeable.
If I randomly assigned treatments, can I generalize to everyone?
No. Random assignment supports a cause-and-effect conclusion about the units in your study. Extending the result to a wider population takes random selection from that population. An experiment run on volunteers still supports causation, but only for units similar to those who took part.