t-test practice problems (one-sample and paired)
By Jude Wallis · Published
This set has 8 problems on one-sample and paired t-tests. Each is a full four-step run: hypotheses and conditions, the standard error and t statistic, degrees of freedom, and a p-value range read from a t-table, then a conclusion in context. Solve each on paper before opening the steps.
AP Statistics: Unit 4 (topics 4.4 Setting Up a Test for a Population Mean or Population Mean Difference, 4.5 Carrying Out a Test for a Population Mean or Population Mean Difference). These problems cover the one-sample t-test for a population mean and the paired t-test for a population mean difference (Unit 4 topics 4.4 and 4.5) in the Fall 2026 AP Statistics course. Unit 4 is 10 to 20% of the multiple-choice section, and the exam supplies the t-table you use to bound each p-value.
What these problems build
These 8 problems build the full workflow of a t-test for a mean: writing the null and alternative hypotheses, checking the random, 10%, and Normal/large-sample conditions, computing the standard error and the standardized test statistic , finding the degrees of freedom, bounding the p-value from a t-table, and stating a conclusion in context. The first problems use a one-sample t-test, which compares a single sample mean (read 'x-bar') against a claimed population mean ('mu-naught'). The later problems use a paired t-test, which works on one column of differences from matched pairs and tests the true mean difference against 0. Both use the same test statistic , where the standard error is with the sample standard deviation and the sample size.
Because the exam hands you a t-table rather than software, each solution bounds the p-value between two table columns instead of reporting one exact number, and technology would give the precise value. Work each problem with the four-step State, Plan, Do, Conclude structure and record the degrees of freedom every time. To review when a mean calls for a t procedure instead of a z procedure, read t test vs z test; to bound a p-value by hand, use the t table; and to check your arithmetic, open the one-sample t test calculator. More sets are on the practice page.
Problem 1
A bottling machine is set to fill sports-drink bottles with a mean of 500 mL. A quality inspector takes a random sample of 16 bottles and finds mL with sample standard deviation mL. At , is there convincing evidence that the machine's true mean fill differs from 500 mL?
Show the worked solution
State. Let be the machine's true mean fill in milliliters. versus , with .
Plan. One-sample t-test for a mean. Random: the 16 bottles are a random sample. 10%: 16 is less than 10% of all bottles the machine fills. Normal/large sample: with , assume the fill amounts are roughly normal with no strong skew or outliers.
Find the standard error. mL.
Find the test statistic and degrees of freedom. , with .
Bound the p-value. This is two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
Conclude. Since the p-value is greater than , which exceeds , fail to reject . There is not convincing evidence that the machine's true mean fill differs from 500 mL.
, , two-sided p-value between and . Fail to reject at ; no convincing evidence the mean fill differs from 500 mL.
Problem 2
A meal-kit company advertises that its dinner recipes take a mean of 30 minutes to cook. A food writer suspects the true mean is longer, times a random sample of 25 recipes, and gets minutes with minutes. At , is there convincing evidence that the mean cook time exceeds 30 minutes?
Show the worked solution
State. Let be the true mean cook time in minutes. versus , with .
Plan. One-sample t-test. Random: the 25 recipes are a random sample. 10%: 25 is less than 10% of all the company's recipes. Normal/large sample: with , assume the cook times are roughly normal with no strong skew or outliers.
Find the standard error. minutes.
Find the test statistic and degrees of freedom. , with .
Bound the p-value. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and .
Conclude. Since the p-value is less than , reject . There is convincing evidence that the true mean cook time exceeds 30 minutes.
, , one-sided p-value between and . Reject at ; convincing evidence the mean cook time exceeds 30 minutes.
Problem 3
A ceramics studio's kiln is supposed to reach a mean peak temperature of 1200°F on its high setting. An artist logs the peak temperature on a random sample of 9 firings and finds °F with °F. At , is there convincing evidence that the kiln's true mean peak temperature differs from 1200°F?
Show the worked solution
State. Let be the kiln's true mean peak temperature in degrees Fahrenheit. versus , with .
Plan. One-sample t-test. Random: the 9 firings are a random sample. 10%: 9 is less than 10% of all firings. Normal/large sample: with , assume the peak temperatures are roughly normal with no strong skew or outliers.
Find the standard error. degrees.
Find the test statistic and degrees of freedom. , with .
Bound the p-value. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence that the kiln's true mean peak temperature differs from 1200°F.
, , two-sided p-value between and . Reject at ; convincing evidence the mean peak temperature differs from 1200°F.
Problem 4
A physical therapist records the increase in grip strength (after minus before), in kilograms, for a random sample of 7 patients who complete a 4-week resistance program: 1, 3, 0, 2, 4, 1, 3. At , is there convincing evidence that the program increases mean grip strength?
Show the worked solution
State. Let be the true mean increase in grip strength (after minus before). versus , with .
Plan. Paired t-test on the 7 differences. Random: the 7 patients are a random sample. 10%: 7 is less than 10% of all such patients. Normal/large sample: with , assume the differences are roughly normal with no strong skew or outliers.
Find the mean difference. kg.
Find the standard deviation of the differences. The squared deviations from 2 are , summing to , so kg.
Find the standard error, test statistic, and df. , so with .
Bound the p-value. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and .
Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence that the program increases mean grip strength.
, , , , one-sided p-value between and . Reject ; convincing evidence the program increases mean grip strength.
Problem 5
A researcher measures systolic blood pressure on the same 6 volunteers with two monitor brands and records the difference (Brand A minus Brand B), in mmHg: 2, -1, 3, -2, 4, 0. At , is there convincing evidence that the two brands give different mean readings?
Show the worked solution
State. Let be the true mean difference in readings (Brand A minus Brand B). versus , with .
Plan. Paired t-test on the 6 differences, one per volunteer measured on both monitors. Random: the problem does not state the 6 volunteers were randomly selected, so treat them as a volunteer (convenience) sample; the conclusion applies to these volunteers and does not generalize to a wider population. Normal/large sample: with , assume the differences are roughly normal with no strong skew or outliers.
Find the mean difference. mmHg.
Find the standard deviation of the differences. The squared deviations from 1 are , summing to , so mmHg.
Find the standard error, test statistic, and df. , so with .
Bound the p-value. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
Conclude. Since the p-value is greater than , which exceeds , fail to reject . There is not convincing evidence that the two monitor brands give different mean readings.
, , , , two-sided p-value between and . Fail to reject ; no convincing evidence the two brands give different mean readings.
Problem 6
A snack company claims its single-serve almond packs contain a mean of 25 g. A consumer group suspects the packs are underfilled, weighs a random sample of 12 packs, and finds g with g. (a) State the hypotheses. (b) Check the conditions. (c) Find the test statistic, degrees of freedom, and p-value range. (d) At , state a conclusion in context.
Show the worked solution
(a) Let be the true mean weight of the packs in grams. Because the group suspects underfilling, versus , with .
(b) One-sample t-test. Random: the 12 packs are a random sample. 10%: 12 is less than 10% of all packs produced. Normal/large sample: with , assume the weights are roughly normal with no strong skew or outliers.
(c) Standard error. g.
(c) Test statistic and degrees of freedom. , with .
(c) p-value. Left-tailed, so use . In the row, falls between (tail ) and (tail ), so the p-value is between and .
(d) Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence that the packs' true mean weight is less than 25 g, supporting the underfilling claim.
(a) , ; (c) , , one-sided p-value between and ; (d) reject , convincing evidence of underfilling below 25 g.
Problem 7
An agronomist pairs 8 plots by location, applies a soil additive to one plot in each pair at random, and records the yield difference (treated minus untreated), in kg: 2, 0, 3, -1, 1, 4, 1, 2. (a) State the hypotheses and check the conditions. (b) Find the test statistic, degrees of freedom, and p-value range. (c) State a conclusion at . (d) Would the conclusion change at ?
Show the worked solution
(a) Let be the true mean yield difference (treated minus untreated). versus , with . This is a matched-pairs design: plots are paired by location and the additive is randomly assigned within each pair, so inference rests on that random assignment. With , assume the differences are roughly normal with no strong skew or outliers.
(b) Mean difference. kg.
(b) Standard deviation of the differences. The squared deviations from 1.5 are , summing to , so kg.
(b) Standard error, test statistic, and df. , so with .
(b) p-value. Two-sided, so use . In the row, falls between (tail ) and (tail ), so one tail holds between and . Doubling, the two-sided p-value is between and .
(c) Conclude at . Since the p-value is less than , which is below , reject . There is convincing evidence that the additive changes mean yield.
(d) At , the p-value is between and , which is greater than , so you fail to reject . The conclusion changes: at the stricter level there is not convincing evidence of a change in mean yield.
, , , , two-sided p-value between and . Reject at , but fail to reject at , so the conclusion changes.
Problem 8
In a sleep study, each of 15 randomly selected participants performs a reaction-time task after a normal night and after a night of restricted sleep. The difference (restricted minus normal), in milliseconds, has mean and standard deviation . (a) State the hypotheses and check the conditions. (b) Find the test statistic, degrees of freedom, and p-value range. (c) State a conclusion at . (d) Interpret the p-value in context and name the error the decision could be.
Show the worked solution
(a) Let be the true mean change in reaction time (restricted minus normal); because researchers expect restricted sleep to slow reactions, versus at . Paired t-test on the 15 differences, each participant measured under both conditions. Random: the 15 participants are a random sample; 10%: 15 is less than 10% of all people the sample represents. Normal/large sample: with , assume the differences are roughly normal with no strong skew or outliers.
(b) Standard error. ms.
(b) Test statistic and degrees of freedom. , with .
(b) p-value. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and .
(c) Conclude. Since the p-value is less than , which is below , reject . There is convincing evidence that restricted sleep increases mean reaction time.
(d) Interpret. If restricted sleep truly had no effect on mean reaction time, there would be between a 1% and 2% chance of getting a sample mean difference at least as large as 12 ms in the slower direction from random variation alone. Because you rejected , the decision could be a Type I error: concluding restricted sleep slows reactions when in truth it makes no difference.
, , one-sided p-value between and . Reject at ; convincing evidence restricted sleep raises mean reaction time. A wrong rejection here would be a Type I error.