Experimental design practice problems (8)
By Jude Wallis · Published
This set has 8 problems on experimental design. They work through four recurring moves: identify the units, variables, and treatments, build or read the design, justify the design choice, then state whether the study supports a cause-and-effect conclusion and who it applies to.
AP Statistics: Unit 1 (topics 1.13 Experimental Design). These problems cover topic 1.13 (Experimental Design), the last topic of Unit 1 in the Fall 2026 AP Statistics course, where Unit 1 carries the heaviest multiple-choice weight at 20 to 30% and describing a design is a standard free-response task.
What these problems build
These 8 problems build the workflow behind every experiment question: naming the experimental units, the explanatory variable with its levels, and the response variable; choosing a design and describing the randomization so a stranger could carry it out; saying what a control group, a placebo, and blinding each accomplish; and deciding what the study licenses you to conclude.
No significance test appears in this set, so there is no test statistic or p-value to compute. The arithmetic instead checks group sizes, counts treatments, and quantifies how lopsided a single random assignment can be, which is the part most students skip. The solutions work through four recurring moves: identify the pieces, build or read the design, justify the choice, state the scope of the conclusion. Problems 1, 2, 5, and 8 run all four; the design write-ups in 3, 4, 6, and 7 stop at the third, because the prompt asks you to build or defend a design rather than to say what it licenses.
A well-designed experiment shows four elements, and a full answer makes all four visible.
- Comparison of at least two treatments, one of which may be a control group.
- [Random assignment](/glossary/random-assignment) of treatments to the experimental units.
- Replication, meaning more than one unit receives each treatment.
- Direct control of other sources of variation, holding them the same from unit to unit.
For the write-up itself, work through how to describe a completely randomized design. More sets are on the practice page.
The three designs and the signal that picks one
The CED names three designs, and the prompt almost always tells you which one it wants.
| Design | The signal in the prompt | How the randomizing runs |
|---|---|---|
| Completely randomized | no extraneous variable is singled out | one pool of units, treatments assigned to all of them at random |
| Randomized block | the prompt names a variable expected to affect the response | sort units into blocks that are alike on that variable, then randomize inside each block |
| Matched pairs | two treatments with units paired, or one unit taking both | randomize within each pair, or randomize the order for each unit |
A matched pairs design is the two-treatment case of a randomized block design, with one block per pair. Blocking is not the same move as stratifying, which happens when you draw a sample rather than when you assign treatments; the two are separated in blocking vs stratifying.
Blocking buys precision. It pulls the variation caused by a known extraneous variable out of the comparison, so the treatment difference is measured against a smaller amount of leftover noise. It does not change the kind of conclusion you may draw.
What a study licenses you to conclude
Two separate questions ride on two separate randomizations, and the last part of most of these problems asks you to keep them apart.
| How the study was run | Random assignment used | No random assignment |
|---|---|---|
| Units randomly selected | cause and effect, generalizes to the population | association only, generalizes to the population |
| Units not randomly selected | cause and effect, these units only | association only, these units only |
Random assignment is what supports causation, because it balances extraneous variables across the treatment groups on average and so blocks confounding. Random selection is what supports generalizing, because it makes the units representative of a population. A study of volunteers with randomly assigned treatments can claim a cause for those volunteers and nothing about the wider population.
When no treatment is assigned, the study is observational, and a confounding variable can always be sitting behind the pattern: something tied to the explanatory variable that also affects the response, whose influence cannot be separated from it. That distinction is worked through in correlation vs causation and confounding vs lurking variable.
Frequently asked questions
Does random assignment guarantee the treatment groups are alike?
No. It balances other variables across the groups on average, over many repetitions, and any single assignment can still come out uneven. Problem 8 puts a number on it: with 2 experienced lifters among 20 volunteers split 10 and 10, both land in the same group with probability . When you already know a variable matters, block on it instead of hoping.
Can an observational study ever support a cause-and-effect conclusion?
No. Without random assignment, any variable tied to the explanatory variable can also be driving the response, and the two effects cannot be separated. Controlling statistically for the variables you thought of does not rule out the ones you did not. The honest report is an association, plus the confounding variables you can name.
Is blocking the same thing as stratifying?
No, and the exam tests the difference. Stratifying happens when you select a sample: you split the population into strata and draw randomly within each. Blocking happens when you assign treatments: you group units that are alike on an extraneous variable and randomize the treatments within each block. Stratifying improves an estimate of a population, blocking sharpens a comparison of treatments.
Does every experiment need a placebo and a control group?
It needs at least two treatments to compare, but the second one can be an active treatment such as the current standard rather than a placebo. A placebo is worth adding whenever the experience of being treated could move the response on its own, which is most of the time with human subjects reporting symptoms or effort.
Problem 1
A gardener tests whether a new LED grow light changes basil growth. She has 36 basil seedlings of the same variety and age. Twelve are randomly assigned to a standard fluorescent light, 12 to the LED at low intensity, and 12 to the LED at high intensity. All 36 get the same pot size, soil, watering schedule, and room temperature. After 4 weeks she records each plant's height in centimeters. (a) Name the experimental units, the explanatory variable, and the response variable. (b) How many treatments are there, and how many units receive each one? (c) Which of the four elements of a well-designed experiment do the shared soil and watering supply? (d) Can a difference in mean height be attributed to the lighting?
Show the worked solution
(a) Identify the pieces. The experimental units are the 36 basil seedlings, since a lighting condition is assigned to each seedling. The explanatory variable is the lighting condition, which the gardener sets. The response variable is plant height in centimeters after 4 weeks.
(b) Count the treatments. The lighting variable has three levels (fluorescent, LED low, LED high), so there are 3 treatments, and seedlings receive each one.
(c) Match the four elements. Comparison: three treatment groups. Random assignment: the lighting conditions are assigned at random. Replication: 12 units per treatment, well above one. Direct control: identical pot size, soil, watering schedule, and room temperature hold extraneous variables fixed from unit to unit, so the shared conditions are the direct control.
(d) Decide on causation. Because the lighting conditions were randomly assigned to the seedlings, other variables are balanced across the three groups on average, so a difference in mean height larger than chance variation would produce can be attributed to the lighting. The seedlings were not randomly selected from a larger population, so the conclusion covers seedlings like these and does not generalize further.
(a) Units: the 36 seedlings; explanatory: lighting condition; response: height in cm after 4 weeks. (b) 3 treatments, seedlings each. (c) Direct control of other sources of variation. (d) Yes, causation is supportable for these seedlings because the lighting was randomly assigned, but the result does not generalize beyond seedlings like these.
Problem 2
A high school compares GPAs across its 480 students. The 180 who play a school sport average a GPA of 3.35, and the 300 who do not average 3.11. The school newspaper reports that playing a sport raises GPA. (a) Name the explanatory and response variables. (b) Is this an experiment or an observational study, and how can you tell? (c) Verify the school's stated overall mean GPA of 3.20 and state the gap between the two groups. (d) Give a confounding variable and explain both links that make it confounding. (e) What would have to change before the newspaper's causal claim is justified?
Show the worked solution
(a) Variables. The explanatory variable is whether a student plays a school sport, a categorical variable with two levels. The response variable is GPA.
(b) Study type. No one assigned students to play a sport or not. Students chose for themselves and the school recorded what had already happened, so this is an observational study. An observational study can establish an association but cannot establish a cause.
(c) Check the overall mean. , which matches the school's figure. The gap between the groups is GPA points.
(d) Name a confounding variable. One plausible confounder is a student's time management and academic motivation, already in place before any tryout. First link: students who organize their time well and care about school are more likely to commit to a season of practices, so the variable is tied to the explanatory variable. Second link: those same habits raise GPA on their own, whether or not the student ever joins a team. The motivated students are concentrated in the sport group, so the effect of playing a sport and the effect of the habits those students already had cannot be separated, and the 0.24 gap could belong to either one or to a mix of both. Note what does not qualify. A study hall that students attend only because they made a team is caused by the explanatory variable rather than sitting behind it, so it is part of how playing a sport might raise GPA, not a rival explanation for the gap. A confounding variable has to be a common cause of both, not a consequence of the treatment.
(e) What would fix it. Only random assignment of students to play or not play would balance motivation and every other pre-existing variable across the two groups on average, so that no one of them systematically favors the sport group. That experiment is neither practical nor ethical here, so the defensible report is that sport participation and GPA are associated at this school, not that one raises the other.
(a) Explanatory: sport participation; response: GPA. (b) Observational study, because students self-selected and no treatment was assigned. (c) , with a gap of GPA points. (d) Time management and academic motivation: students who already have those habits are both more likely to commit to a team and more likely to earn a high GPA, so the two effects on GPA cannot be separated. (e) Random assignment of participation, which is not feasible here, so the paper should report association only.
Problem 3
A pharmaceutical company tests a new allergy tablet on 84 adult volunteers who have seasonal allergies. Describe a completely randomized design that compares the tablet with a placebo, using a symptom score from 0 to 24 recorded after 2 weeks as the response, where a lower score means fewer symptoms. Then say what the placebo and the blinding each accomplish.
Show the worked solution
Number the units. The experimental units are the 84 volunteers. Label them 1 to 84.
Set the group sizes. Two treatments split evenly gives volunteers per group.
Assign at random. Use a random number generator to produce random integers from 1 to 84, ignoring any repeat, until 42 distinct numbers have appeared. Those 42 volunteers take the tablet, and the remaining volunteers take the placebo.
Control everything else. The placebo is a pill identical in size, color, and taste to the tablet but with no active ingredient. Every volunteer takes one pill at the same time of day for the same 2 weeks and receives the same instructions.
Blind both sides. Neither the volunteers nor the nurse who records the symptom scores knows who received which pill, so the study is double-blind.
Measure and compare. After 2 weeks, record each volunteer's symptom score from 0 to 24, then compare the mean symptom score of the 42 tablet takers with the mean score of the 42 placebo takers.
Say what each control buys. The placebo gives a baseline that includes the whole experience of taking a pill, so a difference between the groups points at the active ingredient rather than at the ritual of being treated. The blinding keeps a volunteer's expectations out of the reported symptoms and the nurse's expectations out of the recorded score.
Number the 84 volunteers 1 to 84, generate random integers from 1 to 84 ignoring repeats until 42 distinct numbers appear, give those 42 the tablet and the other 42 an identical placebo, hold dosing time and instructions the same, hide the assignment from both the volunteers and the nurse recording scores, then compare the two groups' mean 0 to 24 symptom scores after 2 weeks. The placebo isolates the active ingredient from the experience of treatment; the double blinding keeps expectations out of the response and the measurement.
Problem 4
A clinic tests a knee-pain cream. Sixty patients are randomly assigned, 30 to the active cream and 30 to an identical inert cream. Patients are not told which cream they received. The physical therapist who scores each patient's knee function at week 6 does know the assignments. (a) Is this study single-blind or double-blind? (b) Name one specific way the therapist's knowledge could distort the results. (c) State the fix. (d) The clinic director asks why the control group could not simply receive nothing. Answer him.
Show the worked solution
(a) Count who is blinded. The subjects do not know which cream they got, but the person measuring the response does. Only one side is hidden, so the study is single-blind, not double-blind.
(b) Name the distortion. A knee function score involves judgment. A therapist who expects the active cream to work can score the same range of motion a point or two higher for the patients he knows received it, pushing the treatment group's mean up for a reason that has nothing to do with the cream.
(c) Fix it. Have the week-6 scoring done by a therapist with no access to the assignment list, with the tubes labeled A and B by someone who measures no one. Then neither the patients nor the person recording the response knows the assignment, which makes the study double-blind.
(d) Explain the inert cream. If the control group received nothing, the two groups would differ in two ways at once: the active ingredient and the experience of applying a cream daily and expecting relief. A patient who expects relief often reports some, and the placebo effect is the difference between the mean response to a placebo and the mean response to no treatment. The inert cream makes the two experiences match on everything except the ingredient under test, so a difference in scores points at the ingredient.
(a) Single-blind, since only the patients are kept unaware. (b) The therapist's expectations can inflate the judgment-based scores he gives to patients he knows received the active cream. (c) Have a therapist with no access to the assignments do the scoring, making the study double-blind. (d) Without the inert cream, the ingredient and the experience of being treated would change together, so the placebo effect would be confounded with the ingredient.
Problem 5
A coach compares two 8-week training programs using 60 runners. Thirty-six of them have finished a marathon before and 24 have not, and prior experience is known to have a large effect on finish time. The response is 10K finish time in minutes at the end of the 8 weeks. (a) Explain why a randomized block design is a better choice here than a completely randomized design. (b) Describe the design, including how the randomization runs and the size of every group. (c) How many runners follow each program in total? (d) Does blocking change what the coach can conclude about cause and effect?
Show the worked solution
(a) Why block. Experience is an extraneous variable with a large effect on finish time. A single completely randomized assignment of all 60 runners could easily leave one program with noticeably more experienced runners, and that imbalance would show up as a program difference. Blocking on experience takes the variation it causes out of the comparison, so the difference between programs is measured more precisely.
(b) Build the blocks and randomize inside each one. Block 1 is the 36 experienced runners and block 2 is the 24 novices. Within block 1, number the runners 1 to 36 and use a random number generator to draw distinct integers from 1 to 36, ignoring repeats, until 18 have appeared; those 18 follow program A and the other follow program B. Within block 2, number the runners 1 to 24 and draw distinct integers from 1 to 24 until 12 have appeared; those 12 follow program A and the other 12 follow program B. Keep coaching hours, surfaces, and the timed 10K course the same for everyone.
(c) Count the four groups. Experienced: runners per program. Novice: runners per program. Each program is followed by runners, and , so every runner is used.
(d) Cause and effect. The treatments are still randomly assigned, now within each block, so a cause-and-effect conclusion is still supported. Compare the two programs within the experienced block and within the novice block, then combine. Blocking buys precision, not a different kind of conclusion, and because the 60 runners were not randomly selected from a population, the finding applies to runners like these.
(a) Experience strongly affects finish time, so blocking on it removes that variation and sharpens the program comparison. (b) Two blocks, randomize inside each: 18 and 18 within the 36 experienced runners, 12 and 12 within the 24 novices. (c) runners per program. (d) No change in kind: random assignment within blocks still supports causation for these runners, and blocking improves precision only.
Problem 6
Eight volunteers each type the same passage on two keyboard layouts, A and B. For each volunteer a coin flip decides which layout is used first, and a 10-minute rest separates the two attempts. Typing speeds in words per minute:
| Volunteer | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Layout A | 62 | 58 | 71 | 49 | 66 | 54 | 60 | 68 |
| Layout B | 65 | 57 | 76 | 54 | 68 | 58 | 63 | 71 |
(a) Name the design and say what the pairing controls. (b) Describe the random assignment and count how many equally likely order assignments there are. (c) Find the eight differences (B minus A) and their mean. (d) What does the coin flip protect against?
Show the worked solution
(a) Name the design. Each volunteer supplies both measurements, so the two treatments are compared within a single person rather than across different people. That is a matched pairs design, the two-treatment case of a randomized block design with one block per volunteer. The pairing controls each volunteer's own typing skill, which varies enormously from person to person and would otherwise swamp any layout difference.
(b) Describe the randomization and count it. For each volunteer, flip a fair coin: heads means layout A is typed first, tails means layout B is typed first. Each of the 8 flips is independent with 2 outcomes, so there are equally likely order assignments, and the study records typing speeds.
(c) Compute the differences. Taking B minus A volunteer by volunteer gives , , , , , , , and . Their sum is , so words per minute in favor of layout B.
(d) Explain the coin flip. If every volunteer typed the same layout first, practice on the passage would raise the second speed and fatigue would lower it, and either effect would push all the differences in one direction. Order would then be confounded with layout, and you could not tell which one produced the gap. Randomizing the order spreads any practice or fatigue effect across both layouts, and the 10-minute rest keeps fatigue small in the first place.
Note the analysis this design calls for. A matched pairs design is analyzed on the single column of 8 differences, not as two independent samples. Testing whether the true mean difference is 0 is the paired t-test worked in t-test practice.
(a) Matched pairs; the pairing controls each volunteer's own typing skill. (b) A coin flip per volunteer sets the order, giving equally likely order assignments across recorded speeds. (c) Differences , summing to 24, so wpm favoring layout B. (d) It keeps practice and fatigue from being confounded with layout.
Problem 7
A greenhouse trial crosses 2 fertilizer brands with 3 watering schedules on 48 identical pots of the same seedling. The greenhouse has a bright south half and a shaded north half, 24 pots in each half, and light has a strong effect on growth. (a) How many treatments does crossing the two variables produce? (b) Set up a randomized block design that uses the greenhouse halves as blocks, and state how many pots receive each treatment within a block and overall. (c) Describe the randomization inside one block. (d) Is light a confounding variable in this design? Explain.
Show the worked solution
(a) Count the treatments. A treatment is one combination of the two explanatory variables, so crossing 2 fertilizer brands with 3 watering schedules gives treatments.
(b) Size the blocks. The blocks are the two greenhouse halves, 24 pots each, because pots within a half receive similar light. Within a block, pots receive each treatment, so each treatment gets pots overall, and the design uses pots, which is every pot available.
(c) Randomize inside a block. Take the south half. Number its 24 pots 1 to 24 and label the 6 treatments 1 to 6. Use a random number generator to draw distinct integers from 1 to 24, ignoring repeats, until 4 have appeared; those pots receive treatment 1. Draw 4 more distinct unused integers for treatment 2, and continue until five treatments hold 4 pots each. The 4 pots whose numbers never came up receive treatment 6. Repeat the whole procedure independently in the north half.
(d) Decide whether light is confounded. It is not. Light is an extraneous variable that clearly affects growth, but every treatment appears exactly 4 times in the bright half and 4 times in the shaded half, so light does not vary with the treatments. Its variation is accounted for by the block. Light would be confounded if, for instance, all the pots on one fertilizer brand had been set in the bright half, because then a growth difference could belong to the brand or to the light and the two effects could not be separated.
(a) treatments. (b) The two halves are blocks of 24 pots; pots per treatment per block and pots per treatment overall, using all 48. (c) Number a half's 24 pots, draw 4 distinct random integers for each of five treatments ignoring repeats, give the 4 pots never drawn the sixth treatment, then repeat in the other half. (d) No. Every treatment appears 4 times in each half, so light varies with the block rather than with the treatments.
Problem 8
A researcher recruits 20 volunteers from a fitness app's mailing list and randomly assigns 10 to a new 6-week strength routine and 10 to the app's standard routine. Two of the 20 have lifted competitively before, and that experience is expected to affect the response, a one-rep-max bench press in kilograms. (a) How many different sets of 10 volunteers could be chosen for the new routine? (b) Find the probability that both experienced lifters end up in the same group. (c) What does that probability say about relying on random assignment alone, and what design change removes the worry? (d) The new routine group averages 4.2 kg more improvement, and the researcher writes that the new routine causes greater strength gains in adults. Critique that claim.
Show the worked solution
(a) Count the possible assignments. Choosing which 10 of the 20 volunteers get the new routine is a combination, and every choice is equally likely: different treatment groups.
(b) Find the probability of a lopsided draw. Condition on the first experienced lifter. Wherever that lifter lands, 9 slots in that group remain and 19 volunteers remain to fill them, all equally likely, so the second experienced lifter joins the first with probability . Counting the assignments directly agrees: both experienced lifters sit in the new-routine group for of the assignments and both sit in the standard group for another , giving favorable assignments and a probability of .
(c) Read what the probability means. Random assignment balances extraneous variables across the groups on average, over many repetitions, not inside any single run. Here there is roughly a 47.4% chance that both experienced lifters land together, which is exactly the imbalance the researcher fears. Blocking removes it: treat the 2 experienced lifters as their own block and randomly send 1 to each routine, then randomly split the remaining 18 volunteers 9 and 9. Every group then holds exactly 1 experienced lifter by design rather than by luck.
(d) Split the claim into its two parts. Causation: the routines were randomly assigned, so a 4.2 kg gap that is larger than chance variation would produce can be attributed to the routine for these 20 volunteers. Generalization: the volunteers came from one app's mailing list and chose to sign up, so they are not a random sample of adults, and the word adults in that sentence is not supported. Rewrite it as: for volunteers like these, the new routine caused a greater mean improvement, and a randomly selected sample of adults would be needed before extending the claim to adults in general.
(a) . (b) , confirmed by counting favorable assignments out of . (c) Random assignment balances on average, not in every single run, so block on experience (1 lifter to each routine, then a random 9 and 9 split of the other 18) to guarantee the balance. (d) If the 4.2 kg gap is larger than chance variation would produce, the causal part is defensible for these 20 volunteers because treatments were randomly assigned, but the leap to adults in general is not, since the volunteers were self-selected from one mailing list rather than randomly selected.