Two-sample t test practice problems and intervals
By Jude Wallis · Published
This set has 8 problems on two-sample t procedures for a difference of means. Each is a full run: hypotheses and conditions on both samples, the standard error, the t statistic with conservative or technology degrees of freedom, then a conclusion in context. Solve each before opening the steps.
AP Statistics: Unit 4 (topics 4.9 Setting Up a Test for the Difference Between Two Population Means, 4.10 Carrying Out a Test for the Difference Between Two Population Means). These problems cover the two-sample t-test for a difference of population means (Unit 4 topics 4.9 and 4.10) in the Fall 2026 AP Statistics course, together with the two-sample t-interval from topics 4.7 and 4.8 that several parts build on.
What these problems build
These 8 problems build the full workflow of a two-sample t procedure: writing hypotheses about , checking conditions on both samples, computing the standard error of the difference, choosing degrees of freedom, bounding a p-value or building an interval, and stating a conclusion in context.
A two-sample procedure uses two independent samples, one from each group, each with its own size , sample mean (read 'x-bar'), and sample standard deviation . The standard error of the difference adds the two variances, never the two standard deviations:
The test statistic and the interval both hang off that one quantity, with ('t-star') the critical value from the t-table:
Reading a t-table by hand gives a bracket rather than a single number, so each solution below bounds the p-value between two of its tail-probability columns; a calculator reports the exact value instead. Either is accepted, as long as you say which one you used. Work each problem with the four-step State, Plan, Do, Conclude structure and record which degrees of freedom you used every time. To check your arithmetic afterward, open the two-sample t test calculator, and to read critical values by hand use the t table. More sets are on the practice page.
Conservative df versus technology df
A two-sample t statistic does not follow an exact t distribution, so its degrees of freedom are approximated. Two approximations are accepted, and you must say which one you used.
| Conservative (by hand) | Technology (calculator) | |
|---|---|---|
| Welch value, usually a decimal | ||
| Effect on p-value | larger | smaller |
| Effect on interval | wider | narrower |
The technology value comes from the Welch formula, which you are never asked to evaluate by hand:
It always lands between and , so the conservative choice never overstates the evidence. Problem 7 is built so the two choices give opposite decisions at , which is exactly why reporting the df matters. For why your calculator prints something like 17.84, read why df is a decimal in a two-sample t test.
Conditions, checked twice
Every condition gets checked once per sample, not once per problem. A response that checks Normality for group 1 and stops loses the point.
- Random. Each group must come from a random sample, or the two groups must come from random assignment in an experiment. Random assignment supports a cause-and-effect claim; random sampling supports generalizing to the population.
- 10%. Only when sampling without replacement from a population: check and separately. Randomized experiments do not need it, because the groups are not samples from a larger population.
- Normal or large sample. For each group, either , or the population is stated to be roughly normal, or a graph of that group shows no strong skew and no outliers.
- Independent groups. The two samples must be independent of each other. If each value in one group is linked to a specific value in the other, the data are matched pairs and this whole procedure is wrong; see paired vs two-sample t test and the t-test practice set for the paired workflow.
Unequal sample sizes are instant proof that a design is not paired. The AP course uses the unpooled two-sample t procedure throughout, so equal population standard deviations are never assumed; see pooled or not.
Frequently asked questions
Which degrees of freedom should I use on the AP exam?
Either is accepted. By hand, use the conservative ; with a calculator, use the Welch value it reports, decimals and all. What you cannot do is compute one and decide with the other. Report the df you used, and make the conclusion follow the p-value that df produced. The conservative choice gives a larger p-value and a wider interval, so it never overstates the evidence.
How do I tell a two-sample t test from a paired t test?
Look at how the data were collected, not at the numbers. If each observation in one group is linked to a specific observation in the other, by the same subject, the same twin, or the same plot, the design is paired and you work with one column of differences. If the groups are separate sets of individuals with no such link, it is a two-sample procedure. Unequal sample sizes rule out pairing immediately.
Do the two populations need equal standard deviations?
No. The AP course uses the unpooled two-sample t procedure, which adds and and never assumes the two population standard deviations match. On a TI-84 that means answering No to the Pooled prompt. Choosing Pooled changes the df to and can produce a different p-value from the one the rubric expects.
Why does my p-value land in a range instead of on one number?
A t-table carries only a fixed set of tail-probability columns, so when you work by hand a t statistic almost always falls between two entries in its df row. Reading the two neighboring columns gives an upper and a lower bound, which is enough to compare against . Technology reports the exact value. A bound that straddles means the evidence is borderline and you should compute the exact p-value before deciding.
Problem 1
A lab tests two brands of AA battery in the same model of wireless mouse. A random sample of 12 Brand A batteries lasts a mean of hours with hours, and an independent random sample of 12 Brand B batteries lasts a mean of hours with hours. At , is there convincing evidence that the two brands differ in true mean lifetime? Use the conservative degrees of freedom.
Show the worked solution
State. Let be the true mean lifetime of Brand A batteries and the true mean lifetime of Brand B batteries, in hours. versus , with .
Plan. Two-sample t-test for a difference of means. Random: both are random samples, selected independently of each other. 10%: 12 is less than 10% of all batteries of each brand. Normal/large sample: , so assume each brand's lifetimes are roughly normal with no strong skew or outliers.
Find the standard error. and , so hours.
Find the test statistic and degrees of freedom. . Conservative .
Bound the p-value. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence that the true mean lifetimes of the two battery brands differ, with Brand A lasting longer in this sample.
, , conservative , two-sided p-value between and . Reject at ; convincing evidence the two brands differ in true mean battery lifetime.
Problem 2
A researcher records daily screen time, in hours. A random sample of 15 students at School A gives and ; an independent random sample of 20 students at School B gives and . Construct and interpret a 95% confidence interval for using the conservative degrees of freedom, and say what it implies about the claim that the two schools have the same true mean screen time.
Show the worked solution
State. Estimate , the difference in true mean daily screen time (School A minus School B), in hours, at 95% confidence.
Plan. Two-sample t-interval for a difference of means. Random: two independent random samples, one at each school. 10%: 15 and 20 are each less than 10% of the students at their own school. Normal/large sample: both samples are under 30, so assume each school's screen times are roughly normal with no strong skew or outliers.
Find the point estimate and standard error. hours. and , so hours.
Find the critical value and margin of error. Conservative . A 95% interval is two-tailed, so read the row in the tail column: . Then hours.
Build the interval. gives hours.
Interpret. We are 95% confident that the interval from to hours captures the true difference in mean daily screen time between School A students and School B students.
Conclude about the claim. Because lies inside the interval, a true difference of zero is a plausible value, so the interval gives no convincing evidence against the claim that the two schools have the same true mean daily screen time.
, , , , . The 95% interval is hours. It contains 0, so there is no convincing evidence of a difference in true mean screen time.
Problem 3
A district compares two review programs. A random sample of 20 students in the online program scores a mean of on a common final with ; an independent random sample of 18 students in the in-person program scores a mean of with . At , is there convincing evidence that the online program has a higher true mean score? Bound the p-value with the conservative degrees of freedom, then compare with technology, which reports and .
Show the worked solution
State. Let be the true mean final score for the online program and the true mean for the in-person program. versus , with .
Plan. Two-sample t-test for a difference of means. Random: two independent random samples, one from each program. 10%: 20 and 18 are each less than 10% of the students enrolled in their own program. Normal/large sample: both samples are under 30, so assume scores in each program are roughly normal with no strong skew or outliers.
Find the standard error. and , so points.
Find the test statistic and conservative degrees of freedom. . Conservative .
Bound the p-value. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and .
Compare with technology. The Welch value gives , which sits inside the same bracket the table produced. The exact conservative p-value is larger than that, because fewer degrees of freedom means heavier tails, and the conservative approach always returns the bigger p-value. Here both fall below , so the decision is the same either way.
Conclude. Since the p-value is less than , reject . There is convincing evidence that the true mean final score is higher for the online program. Because students were sampled from programs they had already joined rather than randomly assigned, this is observational: the higher mean cannot be attributed to the program itself.
, , conservative , one-sided p-value between and (technology: , ). Reject at ; convincing evidence of a higher true mean score for the online program, but no causal claim from an observational study.
Problem 4
Wait times, in minutes, are recorded at two coffee shops. A random sample of 13 customers at Shop A gives and ; an independent random sample of 17 customers at Shop B gives and . (a) Build a 90% confidence interval for using the conservative degrees of freedom. (b) Technology reports , for which . Build that interval too. (c) Say which interval is wider and why, and what both imply about a difference of 0.
Show the worked solution
State. Let and be the true mean wait times, in minutes, at Shop A and Shop B. Estimate at 90% confidence.
Plan. Two-sample t-interval for a difference of means. Random: two independent random samples of customers. 10%: 13 and 17 are each less than 10% of the customers each shop serves. Normal/large sample: both samples are under 30, so assume wait times at each shop are roughly normal with no strong skew or outliers.
(a) Point estimate and standard error. minutes. and , so minutes.
(a) Conservative interval. . A 90% interval leaves in each tail, so the row and the column give . Then , and the interval is minutes.
(b) Technology interval. With and , , and the interval is minutes.
(c) Compare the widths. The conservative interval is minutes wide; the technology interval is minutes wide. Both are centered at the same . The conservative one is wider because describes a t distribution with heavier tails than , so its is larger.
(c) Conclude. Both intervals lie entirely above 0, so 0 is not a plausible value for at 90% confidence either way. We are 90% confident that Shop A's true mean wait is longer than Shop B's by between and minutes on the conservative interval.
. (a) , , interval minutes. (b) , , interval minutes. (c) The conservative interval is wider ( vs minutes) because fewer df means a larger ; both exclude 0, so both show Shop A's true mean wait is longer.
Problem 5
Twelve volunteers are randomly assigned, six to each of two study methods, and then tested on a 20-word list. Words recalled are:
- Method A: 12, 15, 11, 14, 13, 13
- Method B: 10, 13, 9, 11, 8, 9
At , is there convincing evidence of a difference in true mean recall? Use the conservative degrees of freedom.
Show the worked solution
State. Let be the true mean number of words recalled under Method A and the true mean under Method B. versus , with .
Plan. Two-sample t-test for a difference of means. Random: the 12 volunteers were randomly assigned to the two methods, which satisfies the random condition for an experiment. The 10% condition does not apply, because these groups are not samples drawn from a larger population. Normal/large sample: , and because the raw data are given this condition has to be checked on the graphs rather than assumed. Ordered, Method A is and Method B is . A dotplot of each group is single-peaked with no strong skew, and the rule flags no outliers in either group, so a t procedure is reasonable.
Summarize Method A. words. The squared deviations from 13 are , summing to , so words.
Summarize Method B. words. The squared deviations from 10 are , summing to , so words.
Find the standard error, test statistic, and degrees of freedom. and , so words. Then , with conservative .
Bound the p-value. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence of a difference in true mean recall between the two study methods. Random assignment lets you attribute that difference to the method for these 12 volunteers, but because they are volunteers rather than a random sample, the result does not generalize to a wider population.
, , , , , , conservative , two-sided p-value between and . Reject ; convincing evidence of a difference in true mean recall, with Method A higher here. Causal for these volunteers but not generalizable.
Problem 6
A typing tutor is tested by randomly assigning 24 volunteers to two versions of the software. After four weeks, the 13 who used the new version type a mean of words per minute with ; the 11 who used the old version type a mean of with . (a) Explain why a two-sample t procedure fits here rather than a paired one. (b) Check the conditions. (c) At , test whether the new version gives a higher true mean typing speed, using the conservative degrees of freedom. (d) State a conclusion in context and say what the design does and does not let you claim.
Show the worked solution
(a) The two groups contain different people, and no volunteer in one group is linked to a particular volunteer in the other, so the data are two independent samples rather than matched pairs. The unequal group sizes 13 and 11 settle it on their own: a paired procedure needs one difference per pair, which requires equal counts.
(b) Conditions. Random: the 24 volunteers were randomly assigned to the two versions, which is the form the random condition takes in an experiment. Independence: random assignment makes the two groups independent of each other, and no 10% condition is needed because the groups are not samples drawn from a larger population. Normal/large sample: and are both under 30, so assume typing speeds under each version are roughly normal with no strong skew or outliers.
(c) State. Let be the true mean typing speed with the new version and the true mean with the old version, in words per minute. versus , with .
(c) Standard error. and , so words per minute.
(c) Test statistic and degrees of freedom. , with conservative .
(c) p-value. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and .
(d) Conclude. Since the p-value is less than , reject . There is convincing evidence that the new version of the tutor produces a higher true mean typing speed than the old version. Because the version was randomly assigned, a cause-and-effect conclusion about the software is justified for these volunteers. Because they are volunteers rather than a random sample, the result does not generalize to all typists.
(a) Different people in each group and unequal sizes 13 and 11, so two independent samples. (b) Random assignment of the 24 volunteers satisfies the random condition and makes the two groups independent; no 10% condition is needed, because the groups are not samples drawn from a larger population; both , so assume typing speeds under each version are roughly normal with no strong skew or outliers. (c) , , conservative , one-sided p-value between and . (d) Reject ; convincing evidence the new version raises true mean typing speed, causal for these volunteers but not generalizable.
Problem 7
A materials lab measures breaking strength, in kilonewtons, of steel cable from two suppliers. A random sample of 10 lengths from Supplier 1 gives and ; an independent random sample of 10 lengths from Supplier 2 gives and . (a) Test at whether the true mean breaking strengths differ, using the conservative degrees of freedom. (b) Technology reports and a two-sided p-value of . State the decision that follows from it. (c) Explain why the two approaches disagree and what an AP response has to do about it.
Show the worked solution
(a) State. Let and be the true mean breaking strengths, in kilonewtons, of cable from Supplier 1 and Supplier 2. versus , with .
(a) Plan. Two-sample t-test for a difference of means. Random: two independent random samples, one from each supplier's shipment. 10%: 10 lengths is less than 10% of each supplier's production. Normal/large sample: both samples are under 30, so assume breaking strengths from each supplier are roughly normal with no strong skew or outliers.
(a) Standard error and test statistic. and , so kN. Then , with conservative .
(a) Bound the p-value and decide. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and , and the two-sided p-value is between and . Since that exceeds , fail to reject : there is not convincing evidence that the two suppliers differ in true mean breaking strength.
(b) Technology decision. With the two-sided p-value is , which is less than , so reject : there is convincing evidence that the true mean breaking strengths differ.
(c) Why they disagree. The test statistic is identical; only the reference distribution changed. The conservative has heavier tails than , so the same cuts off more area and the p-value grows from past . The conservative method is built to err this way, never claiming more evidence than the data support.
(c) What a response must do. Either method is acceptable, but the response has to report the degrees of freedom it used and state the conclusion that actually follows from that p-value. Mixing a conservative bound with a technology decision earns no credit. With a result this close to , the honest reading is that the evidence is borderline: the samples differ by kN, and larger samples would settle the question.
, . (a) Conservative , two-sided p-value between and , fail to reject . (b) Technology , , reject . (c) Fewer df means heavier tails and a larger p-value; either method is acceptable if you report the df you used and match the conclusion to it.
Problem 8
Two paint formulas are compared for drying time, in minutes. A random sample of 16 boards painted with Formula X dries in a mean of minutes with ; an independent random sample of 21 boards painted with Formula Y dries in a mean of minutes with . Use the conservative degrees of freedom throughout. (a) At , is there convincing evidence that Formula X has a shorter true mean drying time? (b) Construct a 95% confidence interval for . (c) Explain how that interval relates to a two-sided test at . (d) Name the error the decision in (a) could be, and describe it in context.
Show the worked solution
(a) State. Let and be the true mean drying times, in minutes, for Formula X and Formula Y. versus , with .
(a) Plan. Two-sample t-test for a difference of means. Random: two independent random samples of boards. 10%: 16 and 21 are each less than 10% of all boards that could be painted with that formula. Normal/large sample: both samples are under 30, so assume drying times for each formula are roughly normal with no strong skew or outliers.
(a) Standard error. and , so minutes.
(a) Test statistic and degrees of freedom. , with conservative .
(a) Bound the p-value and conclude. Left-tailed, so use . In the row, falls between (tail ) and (tail ), so the p-value is between and . Since the p-value is less than , reject . There is convincing evidence that Formula X has a shorter true mean drying time than Formula Y.
(b) Interval. For 95% confidence with , the column gives , so minutes. The interval is , that is minutes.
(c) Relation to a two-sided test. The whole interval lies below 0, so 0 is not a plausible value for at 95% confidence, and a two-sided test at with would reject . That agrees with part (a): a one-sided p-value between and doubles to a two-sided p-value between and , which is below .
(d) Error. Because was rejected, the decision could be a Type I error: concluding that Formula X dries faster on average when in truth the two formulas have the same true mean drying time. The probability of that error is at most .
, , conservative . (a) One-sided p-value between and ; reject at , convincing evidence Formula X dries faster on average. (b) 95% interval minutes. (c) It excludes 0, matching a two-sided rejection at . (d) A Type I error, with probability at most .