Difference of two sample means practice
By Jude Wallis · Published
This set has 8 problems on the sampling distribution of the difference between two sample means. You find its center, add the two variances to get its standard deviation, decide when a normal model is justified, and compute probabilities about a difference. Solve each before opening the steps.
AP Statistics: Unit 4 (topics 4.6 Sampling Distributions for the Difference Between Two Sample Means). These problems cover Unit 4 topic 4.6 of the Fall 2026 AP Statistics course: the mean, the standard deviation, and the shape of the sampling distribution of the difference between two sample means, plus probabilities computed from it. Every problem gives population standard deviations, so the root produces a true standard deviation and the standardized value is a z rather than a t, with version B of problem 8 the one place sample standard deviations appear. A known sigma settles which reference curve, not whether a normal curve applies: shape is a separate condition, cleared by a normal population or by a sample that reaches 30, and problem 4 gives both population standard deviations and still fails it. Once only sample standard deviations are available, the same expression becomes the standard error behind the two-sample t procedures in topics 4.7 through 4.10. Unit 4 carries 10 to 20 percent of the multiple-choice section.
What these problems build
Topic 4.6 asks three questions about (read 'x-bar-one minus x-bar-two'), the difference between two independent sample means: where it is centered, how much it varies, and when a normal curve describes it.
- Center. The difference of the sample means sits on the difference of the population means, at every pair of sample sizes and for every pair of population shapes.
- Spread. The two quantities under the root, and , are the variances of and . Add them, then take the square root once, at the end.
- Shape. Checked once per sample, not once per problem. Each sample clears on one of two routes: its population is normal, which works at any , or its own sample size reaches 30. Both samples have to clear, and one failure closes the normal model for the difference.
Problems 1 through 3 are built on the spread formula, problem 4 on the shape argument, problems 5 and 6 on probabilities, problem 7 on choosing a sample size, and problem 8 on the one distinction that decides whether any of this is a calculation at all. For the concept behind all of it read sampling distributions explained and the central limit theorem guide, for the one-sample version work the x-bar set first, and for the topic page see 4.6. Check your areas against the normal distribution calculator or the z-table, and review how to find a z score if standardizing is shaky.
Why the variances add when you subtract
This is the whole set, and it is the most common error on the topic. Standard deviations never add. Variances do, for independent quantities, and they add for a difference exactly as they do for a sum. Problem 1 uses numbers that make the point visible: the two sample means have standard deviations 4 and 3.
| Method | Problem 1's numbers | Result |
|---|---|---|
| Add the variances, root at the end (correct) | ||
| Add the standard deviations | , 40% too large | |
| Subtract the variances | , far too small |
Variance measures squared distance from the center, so two independent sources of variability combine the way the legs of a right triangle do: 4 and 3 give 5, not 7. Adding the standard deviations always overstates the spread, which makes an observed gap look ordinary when it is not.
Subtracting is the error that feels most natural and lands furthest from right. Subtracting the means cancels nothing. Sample 1 misses its target by some amount, sample 2 misses its target by some amount, and both misses land in the gap between them, so the gap wobbles more than either mean alone does. Watch what 'subtract the variances' predicts when the two terms happen to be equal: a standard deviation of zero, meaning every pair of samples returns exactly the same difference. Problem 3 sets Cleo's version of that formula beside a simulation, and its solution works through the collapse. For why variance is the quantity that behaves this way, read why variance is in squared units, and note that the same sum reappears as the standard error of every two-sample t procedure.
Sigma or s, and why it decides the procedure
The other classic error on this topic is running a calculation on numbers that only support a procedure. The arithmetic is identical either way, which is exactly what makes the mistake easy to miss.
| Topic 4.6, this set | Topics 4.7 through 4.10 | |
|---|---|---|
| Spread figures given | population and | sample and |
| What the root produces | a standard deviation | a standard error |
| Reference curve | standard normal, | distribution |
| Question asked | a probability about | a confidence interval or a test |
Every problem in this set hands you population standard deviations, so the root produces a genuine standard deviation and the standardized value is a . Real two-sample data almost never comes that way: you estimate the spread from the samples themselves, that estimate carries its own sampling error, and the reference curve widens into a . Problem 8 runs one study both ways so the difference is concrete, and its version B, the one place in this set where sample standard deviations appear, is the only below.
Keep the two conditions apart, because the Greek letter settles only one of them. Knowing tells you the spread, and the spread is what decides which curve you standardize against. Whether a normal curve applies at all is the shape condition from the first section, and it is cleared only by a normal population or by a sample that reaches 30. Those two checks are independent, and both have to pass before a normal-curve area exists. Problem 4 hands you and and still refuses the normal model, because its second sample is 20 observations from a skewed population.
See t test vs z test, standard error vs standard deviation, and z-interval vs t-interval. More sets are on the practice page.
Frequently asked questions
Why do the variances add when I am subtracting the means?
Because both samples miss their targets, and neither miss cancels the other. For independent quantities the variances add whether you subtract or combine any other way, so and the square root comes last. Adding the two standard deviations instead always gives a number that is too large, as problems 1 and 2 quantify.
Do both samples need to reach 30, or just one?
Both, when the populations are not normal. The condition is checked once per sample: each one qualifies either because its population is stated to be normal, which works at any size, or because its own reaches 30. A sample of 45 from a skewed population does not cover a partner sample of 20 from another skewed population. Problem 4 is built on exactly that case.
Does it matter which group I call group 1?
Only for the sign. Swapping the labels flips the center from to , but the standard deviation is identical, because swapping just exchanges two terms that are being added. Pick an order, state it in words, and keep it. A question about group 2 finishing higher is a question about a negative difference, which is what problem 6 works through.
Can the standard deviation of the difference ever be smaller than one group's own?
No. Since is positive, the sum under the root is strictly larger than alone, so the difference always varies more than either sample mean does. That is a hard floor: enlarging one group can push the standard deviation of the difference down toward the other group's value but never past it. Problem 7 uses it.
When is this a z problem and when is it a t problem?
Look at which letters the problem hands you. Population standard deviations and give a genuine standard deviation for the difference, so you standardize to a ; that is topic 4.6. Sample standard deviations and give a standard error that carries its own sampling error, so the reference curve is a and you are in topics 4.7 through 4.10. The arithmetic under the root is identical, which is what makes the mixup easy. The letters decide the curve and nothing else: a normal model still has to be earned separately, once per sample, which is what problem 4 turns on. See t test vs z test.
Problem 1
Two coaching programs report scores on a common 100-point exam. Scores are approximately normal in both populations. Program 1 has points and points; program 2 has points and points. A researcher takes independent random samples of students from program 1 and from program 2. (a) Find the mean and the standard deviation of the sampling distribution of . (b) A student finds that has standard deviation 4 and has standard deviation 3, then reports . Say what is wrong and by how much the answer is off.
Show the worked solution
(a) Center. points. The difference of the sample means is centered on the difference of the population means, so this is just a subtraction.
(a) First variance term. . This is the variance of , and its square root, points, is the standard deviation of .
(a) Second variance term. , so has standard deviation points.
(a) Add the variances, then take the root once. points.
(b) Name the error. The numbers 4 and 3 are standard deviations, and standard deviations do not add. Their squares do: , and the root of 25 is 5, not 7. The student's 7 is times the correct value, so it overstates the spread by 40%.
(b) Why squares. Variance is built from squared distances, so two independent sources of variability combine like the legs of a right triangle: legs of 4 and 3 give a hypotenuse of 5. Adding standard deviations would only be right if the two samples moved in perfect lockstep, and independent samples do not.
Note the direction of the mistake. An inflated standard deviation makes any observed gap look closer to the center than it really is, so it hides real differences rather than inventing them.
(a) points and points. (b) The student added standard deviations; only variances add. gives 5 points, so 7 overstates the spread by 40%.
Problem 2
Commute times are right-skewed in two cities. City A has minutes and minutes; city B has minutes and minutes. Independent random samples of commuters and commuters are taken. (a) Justify a normal model for and find its mean and standard deviation. (b) Find . (c) An analyst instead computes the standard deviation of each sample mean and adds those two numbers. Report the standard deviation and the probability that mistake produces, and say which way the error pushes the conclusion.
Show the worked solution
(a) Shape, checked once per sample. Both populations are right-skewed, so the normal-population route is closed for both. Sample A has and sample B has , so the central limit theorem covers each sample mean, and the difference is approximately normal. Both had to clear 30; one sample at 40 would not rescue the other at 12.
(a) Center. minutes.
(a) Spread. and , so minutes.
(b) Standardize and find the area. , and . Rounding to 1.20 for a table gives 0.1151, close enough to pick the same multiple-choice option.
(c) Redo it the wrong way. The two sample means have standard deviations and minutes. Adding those gives 2.3604 minutes, which is times the correct 1.6733, about 41% too large.
(c) The damaged probability. , and , about 1.7 times the correct 0.1160.
(c) Direction. Adding standard deviations always inflates the spread, so a gap of 5 minutes looks far more ordinary than it is. Carried into inference, the same inflation widens confidence intervals and enlarges p-values, so the mistake pushes every conclusion toward 'no difference found'.
(a) Approximately normal by the central limit theorem, since and both reach 30; mean 3 minutes and standard deviation minutes. (b) , so . (c) Adding the two standard deviations gives 2.3604 minutes, about 41% too large, and , about 1.7 times the correct value. The error always overstates the spread and understates how unusual a gap is.
Problem 3
A class simulates a difference of sample means. Population 1 is normal with and ; population 2 is normal with and . Each repetition draws independent random samples of and and records . Across 10,000 repetitions the simulated differences have mean 6.04 and standard deviation 3.58. Three students propose a formula for the standard deviation: Ana uses , Ben uses , and Cleo uses . (a) Evaluate all three. (b) Say which one the simulation supports and why the other two fail. (c) Explain why the simulated mean is 6.04 rather than exactly 6.
Show the worked solution
Set up the two building blocks. and . Their square roots, 2 and 3, are the standard deviations of and .
(a) Ana. .
(a) Ben. .
(a) Cleo. .
(b) Compare with the simulation. The simulated standard deviation is 3.58, which matches Ben's 3.6056 and nothing else. Ana's 5 is 39% too large and Cleo's 2.2361 is 38% too small, both far outside what 10,000 repetitions would miss by.
(b) Why Ana fails. She added standard deviations. Only variances add for independent quantities, so the root has to come after the addition, not before.
(b) Why Cleo fails. She read 'difference of means' as 'difference of variances'. Subtracting the sample means does not remove either sample's variability, it combines both. Her formula also collapses in two obvious ways: it returns 0 whenever the two terms are equal, predicting a difference that never varies at all, and it asks for the root of a negative number whenever the first term is larger.
(c) Simulation error. A simulated mean is itself a statistic computed from 10,000 draws, so it estimates the true center rather than landing on it. The exact values are and ; the simulation returned 6.04 and 3.58, and more repetitions would tighten both toward those numbers.
(a) Ana 5, Ben , Cleo . (b) The simulated 3.58 supports Ben. Ana added standard deviations instead of variances, which always overstates. Cleo subtracted the variances, which understates, returns 0 when the two terms match, and goes imaginary when the first is larger. (c) 6.04 and 3.58 are estimates from 10,000 repetitions; the exact values are 6 and 3.6056.
Problem 4
Daily app-use times are strongly right-skewed in two user groups. Group 1 has minutes and minutes; group 2 has minutes and minutes. Independent random samples of and users are taken. (a) Find the mean and standard deviation of . (b) Can you compute ? Justify, and say which of your part (a) answers survive your verdict. (c) Suppose instead that group 2's use times were stated to be approximately normal, with everything else unchanged. Answer (b) again.
Show the worked solution
(a) Center. minutes.
(a) Spread. and , so minutes. Note that the smaller sample contributes the larger variance term even though its population is less variable, because sits in the denominator.
(b) Check each sample separately. Sample 1: the population is skewed, so the normal-population route is closed, but , so the central limit theorem covers it. Sample 2: the population is skewed and , so neither route applies.
(b) Verdict. One failure is enough. Sample 2 clears no route, so cannot be modeled as normal, so neither can the difference, and no normal-curve probability may be computed. The answer is no.
(b) What survives. Both answers in part (a) stand. Neither the center nor the standard deviation depends on the shape of a population, so minutes and minutes are still exactly right. Only the area under a normal curve becomes unavailable.
(c) Recheck the shape. A normal population qualifies its sample at any size, so with group 2 normal, sample 2 clears at . Sample 1 still clears through the central limit theorem at . Both samples clear, so the difference is approximately normal, with the same center 7 and the same standard deviation 5.3852 minutes.
(c) Standardize and find the area. , and .
(a) minutes and minutes. (b) No: group 2's population is skewed and , so that sample clears neither route to normality; the center and standard deviation from (a) are unaffected, since neither depends on shape. (c) Yes: a normal population qualifies sample 2 at any size, so and .
Problem 5
Mature heights of two tomato varieties are approximately normal. Variety A has cm and cm; variety B has cm and cm. A grower measures independent random samples of variety A plants and variety B plants. (a) Describe the sampling distribution of , including why the sample sizes are not a problem. (b) Find . (c) The grower doubles both sample sizes, to 24 and 32. Recompute the probability and explain what changed.
Show the worked solution
(a) Shape. Both populations are stated to be approximately normal, so each sample mean is approximately normal at any sample size and so is their difference. The central limit theorem is not needed, which is why raises no objection here even though it is well under 30.
(a) Center and spread. cm. and , so cm.
(b) Standardize and find the area. , and , roughly 1 sample pair in 15.
(c) Recompute the spread. and , so the two variance terms are each exactly half what they were, their sum falls from 7 to 3.5, and cm. Doubling both sample sizes halved the total variance, so the standard deviation fell by a factor of , from 2.6458 to 1.8708.
(c) Recompute the probability. , and , roughly 1 sample pair in 61.
(c) Read it back. The true gap between the varieties is still 4 cm; larger samples never move the center. What changed is precision: a sampled gap of 8 cm now sits 2.14 standard deviations from the center instead of 1.51, so it fell from about a 1 in 15 result to about a 1 in 61 result.
(a) Approximately normal because both populations are normal, so no sample-size threshold applies; mean 4 cm and standard deviation cm. (b) , so . (c) cm, , and . Doubling both sample sizes halved the total variance and divided the standard deviation by , leaving the center at 4 cm.
Problem 6
Two job-training programs are compared on an end-of-course test. Scores are somewhat skewed in both populations. Program 1 has points and points; program 2 has points and points. Independent random samples of 36 graduates are taken from each program. (a) Find the mean and standard deviation of , justifying the normal model. (b) Find the probability that program 2's sample mean comes out higher than program 1's. (c) Interpret that number for someone about to run this study.
Show the worked solution
(a) Shape. Both populations are skewed, so the normal-population route is closed for both, but , so the central limit theorem covers each sample mean and the difference is approximately normal.
(a) Center and spread. points. and , so points.
(b) Translate the event before computing anything. 'Program 2's sample mean is higher' means , which is the same as . Written that way it is an ordinary left-tail question about the distribution from part (a).
(b) Standardize and find the area. , and .
(c) Interpret. Program 1 really is 3 points better on average, yet about 11.5% of studies built this way, roughly one in nine, would report program 2 ahead. Nothing would have gone wrong in those studies; sampling variability alone reverses the ranking that often.
(c) Say what drives it. The true gap of 3 points is only standard deviations above zero, so zero is not far out in this distribution. Larger samples shrink and push the reversal probability down; nothing else in the design will.
(a) Approximately normal by the central limit theorem, since both samples reach 36; mean 3 points and standard deviation points. (b) The event is , so and . (c) Even though program 1 is truly 3 points better, about 11.5% of studies of this size, roughly one in nine, would show program 2 ahead, because the true gap is only 1.2 standard deviations from zero.
Problem 7
A greenhouse study compares plant growth under two treatments. Treatment A has already been run on plants, and long-run records give cm. Treatment B has cm, and the researcher can still choose . (a) Find the smallest for which the standard deviation of is at most 2.5 cm. (b) Is there any that brings it to 1.9 cm or below? Justify your answer.
Show the worked solution
Identify what is fixed. Treatment A is done, so is locked in. Only the second variance term, , is still under the researcher's control.
(a) Set up the inequality. Require . Square both sides: , so .
(a) Solve. , so the smallest whole number that works is .
(a) Check both sides of the cutoff. At , exactly, which satisfies 'at most 2.5'. At , , so 35 fails.
(b) Find the floor. is positive for every sample size, so and therefore cm no matter how large becomes. Since , no works.
(b) State the general lesson. The standard deviation of a difference can never fall below the standard deviation of either sample mean on its own, here cm. Pouring data into one group has diminishing returns and then stops helping; getting under 2 cm requires enlarging sample A.
(a) , so and the smallest value is , which gives exactly 2.5 cm ( gives 2.5128 cm). (b) No. cm for every , and , so only a larger sample from treatment A can push the standard deviation below 2 cm.
Problem 8
A dietitian compares sodium per serving in two soup brands, and sodium content is approximately normal for both. Independent random samples of cans of brand 1 and cans of brand 2 are analyzed. Assume the two brands truly have the same mean sodium content. (a) Version A: the manufacturers' long-run process records give mg and mg. Find the standard deviation of and the probability that the two sample means differ by more than 30 mg in either direction. (b) Version B: no process records exist, and the only spread figures are the sample standard deviations from these very cans, mg and mg. State what changes and what does not.
Show the worked solution
(a) Center. The brands are assumed equal, so mg.
(a) Spread. and , so mg.
(a) Shape. Both populations are approximately normal, so the difference is approximately normal at these sizes and needs no appeal to the central limit theorem.
(a) Probability. 'Differ by more than 30 mg in either direction' is a two-tailed event. , so .
(b) What does not change. The arithmetic. mg either way, because the expression is the same one. The center is still 0 mg under the assumption of equal means, and the variances still add.
(b) What does change. In version A, 14.5430 mg is a known standard deviation, so the standardized value is a and its tail area comes from the standard normal curve. In version B it is a standard error: an estimate built from the same 55 cans, carrying its own sampling error. The standardized value is then a , compared against a distribution rather than the standard normal.
(b) Size of the difference it makes. Referring the same statistic 2.063 to a distribution with the unpooled (Welch) degrees of freedom, , gives a two-sided value of 0.0451 rather than 0.0391. The curve has heavier tails, so it always returns the larger area.
(b) The rule to carry. A Greek decides which curve you standardize against, not whether a normal model exists. With known the root is a true standard deviation, so once the shape condition is met the reference curve is the standard normal: that is topic 4.6. A Roman computed from the data makes the root a standard error and the reference curve a : that is topics 4.7 through 4.10. This problem carries one letter in each version, which is why it can be asked both ways. Problem 4 is the reminder that the shape condition is a separate check, and that it can fail while is known.
(a) mg, , and . (b) The number 14.5430 mg is unchanged, but it becomes a standard error rather than a known standard deviation, the standardized value follows a distribution instead of the standard normal (two-sided 0.0451 at instead of 0.0391), and the question moves from a topic 4.6 probability to the two-sample t procedures of topics 4.7 through 4.10.