Inference for means: 8 mixed t procedure problems
By Jude Wallis · Published
This set has 8 problems spanning Unit 4 topics 4.2 to 4.10. Nothing tells you which procedure to run, so every problem starts with the same decision: one-sample t interval, one-sample t test, two-sample t, or paired t. Several hinge entirely on reading the design rather than on arithmetic.
AP Statistics: Unit 4 (topics 4.2 Constructing a Confidence Interval for a Population Mean or Population Mean Difference, 4.4 Setting Up a Test for a Population Mean or Population Mean Difference, 4.5 Carrying Out a Test for a Population Mean or Population Mean Difference, 4.7 Constructing a Confidence Interval for the Difference Between Two Population Means, 4.9 Setting Up a Test for the Difference Between Two Population Means, 4.10 Carrying Out a Test for the Difference Between Two Population Means). This set spans Unit 4 topics 4.2 through 4.10 of the Fall 2026 AP Statistics course, the one-sample and two-sample means procedures together. Unit 4 is 10 to 20% of the multiple-choice section, and free-response questions on it routinely make the choice of procedure a scored component rather than a given.
What these problems build
Individual practice sets tell you the procedure in their title. The exam does not. It hands you a paragraph about how data were collected and expects you to name the procedure, justify it, check its conditions, run it, and interpret the result in context. Picking the wrong procedure loses every point downstream, no matter how clean the arithmetic is.
These 8 problems are deliberately unlabeled. Problem 1 is identification only and problem 7 works backwards from a printed interval; the other six name the procedure first and then carry it through to a conclusion. Two of them are built around a design that looks like something it is not.
Before starting, review the conditions each procedure needs in the conditions for inference checklist, and if a scenario is genuinely ambiguous, work it through the which test decision tool. For arithmetic checks, use the one-sample t test calculator or the two-sample t test calculator. More sets are on the practice page.
Three questions that settle the procedure
Ask them in this order, and the answer falls out.
1. How many groups of data are there? One group compared against a claimed number is a one-sample procedure. Two groups is either paired or two-sample, which question 2 settles.
2. If there are two columns, is each value in one tied to a specific value in the other? Same subject measured twice, same specimen tested twice, or two subjects matched and then split: paired, so build one column of differences and run a one-sample procedure with . Separate sets of individuals with no such link: two-sample, with by hand.
3. Is the question asking how big, or asking whether? 'Estimate', 'construct an interval', and 'with 95% confidence' call for an interval. 'Is there convincing evidence', 'test the claim', and a stated call for a test.
| Design | Parameter | Standard error | df by hand |
|---|---|---|---|
| One sample vs a claim | |||
| Paired or matched | |||
| Two independent samples |
Two traps recur. Equal sample sizes do not make a design paired: two groups of 15 different people are still two independent samples. And comparing one sample against a fixed published number is a one-sample procedure, not a two-sample one, because a number is not a sample. See which statistical test to use and paired vs two-sample t test.
Frequently asked questions
What is the fastest way to tell the three t procedures apart?
Count the groups first. One group measured against a claimed value is one-sample. Two columns of numbers where each value in one is tied to a specific value in the other, by the same subject, the same specimen, or a deliberate match, is paired: collapse to differences and run the one-sample procedure. Two columns from separate individuals with no such link is two-sample. Unequal sample sizes settle it immediately, since pairing requires equal counts.
Is comparing my sample to a published number a two-sample test?
No. A two-sample procedure needs two sets of data so that both contribute sampling variability to the standard error. A published guideline, a labeled weight, or an advertised claim is a single fixed number with no sample behind it, so it goes in as in a one-sample test. Problem 6 works through exactly this mistake.
How do I know whether to build an interval or run a test?
Read the verb. 'Estimate', 'construct an interval', and any stated confidence level ask how big the parameter is, which is an interval. 'Is there convincing evidence', 'test the claim', and a stated significance level ask whether a specific value is plausible, which is a test. They are two views of the same information: a interval and a two-sided test at level always give the same decision about a value inside or outside the interval. See confidence interval vs hypothesis test.
Does picking the wrong procedure cost every point on an FRQ?
Close to it. The identification and justification usually carry their own scoring component, and the computation is scored against the procedure the design actually calls for, so correct arithmetic on the wrong procedure earns little. The cheapest insurance is one sentence naming the procedure and the design feature that forces it, written before any calculation: 'paired test, because each subject was measured under both conditions.'
Problem 1
Name the procedure each study calls for, and give the reason. Do not compute anything.
(a) A random sample of 40 used cars is tested, and the goal is to estimate the mean carbon dioxide output per kilometer with 95% confidence. The population standard deviation is unknown. (b) Twenty-five volunteers each taste two formulas of the same soda and rate both on a 100-point scale. (c) An independent random sample of 30 seniors and an independent random sample of 30 first-year students report nightly sleep, and the question is whether mean sleep differs. (d) A random sample of 18 batteries is tested against the manufacturer's claim that mean life is 500 hours.
Show the worked solution
(a) One-sample interval for . There is one group, no claimed value to test against, and the word 'estimate' with a confidence level. The interval is rather than because is unknown and stands in for it, with .
(b) Paired procedure on the differences. Each volunteer produces two ratings, so volunteer 12's two scores belong together. Collapse to 25 differences and run a one-sample procedure with . Whether it is an interval or a test depends on what is asked next, but the pairing is fixed by the design.
(c) Two-sample test for . Two separate groups of students with no link between any senior and any first-year, and the question asks whether means differ, which is a test rather than an interval. By hand, . Equal group sizes here are a coincidence, not pairing.
(d) One-sample test for . One group measured against a fixed claimed value, hours, with . The 500 is a number, not a second sample, so no two-sample procedure is involved.
Notice the pattern. Nothing in the arithmetic decided any of these. The number of groups, the presence or absence of a link between the columns, and the verb in the question did all the work.
(a) One-sample interval, . (b) Paired procedure on 25 differences, . (c) Two-sample test, by hand. (d) One-sample test against , .
Problem 2
A 3D printer's manufacturer advertises a mean print time of 45 minutes for one benchmark model. A maker space times a random sample of 20 prints of that model on its own machine and finds minutes with minutes; a dotplot of the 20 times is roughly symmetric with no outliers. At , is there convincing evidence that this machine's true mean print time differs from the advertised 45 minutes?
Show the worked solution
Choose the procedure. One group of 20 prints compared against a fixed advertised value, and is unknown, so this is a one-sample test for a mean. There is no second sample to compare with.
State. Let be the true mean print time, in minutes, for this benchmark model on this machine. versus , with . The alternative is two-sided because 'differs' names no direction.
Plan. Random: the 20 prints are a random sample. 10%: 20 is less than 10% of all prints this machine could produce. Normal/large sample: , so lean on the dotplot, which is roughly symmetric with no outliers.
Compute the standard error and the test statistic. minutes, so with .
Find the p-value. Two-sided, so use . In the row it falls between (tail ) and (tail ), so one tail holds between and and the two-sided p-value is between and . Technology gives .
Conclude. The p-value is below , so reject . There is convincing evidence that this machine's true mean print time differs from the advertised 45 minutes, running longer in this sample. The margin is thin: at the same data would not have rejected, so the evidence is real but not overwhelming.
One-sample test. , , , two-sided p-value (table bracket to ). Reject at ; convincing evidence the true mean print time differs from 45 minutes, though the result would not reject at .
Problem 3
A biologist measures forewing length, in millimeters, on an independent random sample of 15 butterflies from Meadow Site and an independent random sample of 15 from Woodland Site. Meadow gives with ; Woodland gives with . A student says that because both samples have 15 observations, the data can be paired. (a) Explain why that is wrong and name the correct procedure. (b) Carry it out at using the conservative degrees of freedom. (c) State what the study design does and does not let the biologist claim.
Show the worked solution
(a) Why pairing fails. Pairing requires a link between one specific observation in each group, created by measuring the same individual twice or by matching individuals deliberately. Here the two samples are different butterflies caught at different sites, chosen independently, so butterfly 7 at Meadow has no partner at Woodland. Equal sample sizes are a necessary consequence of pairing, never a cause of it. The correct procedure is a two-sample test for a difference of means.
(b) State. Let be the true mean forewing length at Meadow Site and the true mean at Woodland Site, in millimeters. versus , with .
(b) Plan. Two independent random samples, one from each site. 10%: 15 is less than 10% of the butterflies at each site. Normal/large sample: both samples are under 30, so assume forewing lengths at each site are roughly normal with no strong skew or outliers. Independence: the samples were drawn independently of each other.
(b) Standard error. and , so mm. Note that variances are added, never standard deviations.
(b) Test statistic and p-value. , with conservative . Two-sided, so use , which falls between (tail ) and (tail ), putting the two-sided p-value between and . Technology gives .
(b) Conclude. The p-value is below , so reject . There is convincing evidence that true mean forewing length differs between the two sites, with Woodland longer in these samples.
(c) Scope. Both samples were randomly selected, so the conclusion generalizes to the butterflies at each of these two sites. Site was not assigned by the biologist, so this is observational: the difference cannot be attributed to the habitat itself, since anything that differs between the sites, food supply, predation, or which subspecies happens to live there, could produce the same gap.
(a) Different butterflies with no link between individuals, so the data are two independent samples; equal sizes do not create pairing. (b) mm, , conservative , two-sided p-value (table bracket to ). Reject ; convincing evidence the true mean forewing lengths differ. (c) Random selection generalizes to each site, but with no random assignment the difference cannot be attributed to habitat.
Problem 4
An environmental lab checks two chlorine test kits by running both kits on each of 12 randomly selected water samples. For the differences (Kit 1 reading minus Kit 2 reading), in parts per million, the lab reports and . (a) Name the procedure and say what feature of the design fixes it. (b) Construct a 95% confidence interval for the true mean difference. (c) Interpret the interval and say what it implies about whether the kits agree.
Show the worked solution
(a) Identify the design. Each water sample is measured by both kits, so the two readings on sample 5 belong together and their difference is meaningful. That link makes this a matched-pairs design, and the correct procedure is a paired interval, which is a one-sample interval built on the 12 differences with . There are 24 readings but only 12 pairs, and .
(b) Plan. Random: the 12 water samples are a random sample. 10%: 12 pairs is under 10% of the water samples the lab could test. Normal/large sample: , so assume the differences are roughly normal with no strong skew or outliers. Conditions are checked on the differences, not on the two columns of readings.
(b) Standard error. ppm.
(b) Critical value and margin of error. With , the 95% column gives , so ppm.
(b) Interval. ppm.
(c) Interpret. We are 95% confident that the interval from to ppm captures the true mean difference in chlorine reading (Kit 1 minus Kit 2) on water samples like these.
(c) Answer the agreement question. The interval lies entirely above 0, so a true mean difference of zero is not plausible: the kits do not agree on average, and Kit 1 reads higher. A matching two-sided test confirms it, with at giving a p-value of , below . Note the practical caveat: the lower bound is only ppm, so the bias might be too small to matter for the lab's purposes even though it is statistically detectable.
(a) Paired interval, because both kits measure the same 12 water samples, so pairs and . (b) ppm, , , interval ppm. (c) 95% confident it captures the true mean difference; it excludes 0, so Kit 1 reads higher on average, though the difference may be too small to matter in practice.
Problem 5
A poultry researcher compares two feed formulas. An independent random sample of 14 hens on Feed A lays a mean of 232 eggs per year with ; an independent random sample of 16 hens on Feed B lays a mean of 221 eggs with . Construct a 90% confidence interval for the difference in true mean annual egg production using the conservative degrees of freedom, and interpret it.
Show the worked solution
Choose the procedure. Two independent random samples of different hens, and the question asks for an interval at a stated confidence level, so this is a two-sample interval for . The unequal sizes 14 and 16 rule out any paired procedure on sight.
State. Let and be the true mean annual egg production for hens on Feed A and Feed B. Estimate with 90% confidence.
Plan. Random: two independent random samples, one per feed. 10%: 14 and 16 are each less than 10% of the hens that could be raised on that feed. Normal/large sample: both samples are under 30, so assume egg counts under each feed are roughly normal with no strong skew or outliers.
Point estimate and standard error. eggs. Then and , so eggs.
Critical value and margin of error. Conservative . A 90% interval leaves in each tail, so the row gives , and eggs.
Build the interval. eggs per year.
Interpret. We are 90% confident that the interval from to eggs per year captures the true difference in mean annual egg production between hens on Feed A and hens on Feed B.
Read the zero. The interval contains 0, so a true difference of zero is a plausible value and this study gives no convincing evidence that the two feeds differ in true mean production. The interval is nearly 26 eggs wide, so 'no convincing evidence' here means the study is too small to settle the question, not that the feeds are known to perform identically.
Two-sample interval. eggs, , conservative , , . The 90% interval is eggs per year. It contains 0, so there is no convincing evidence of a difference, and the interval is too wide to settle the question either way.
Problem 6
A public health guideline says a single frozen dinner should average no more than 600 calories. A researcher measures a random sample of 22 frozen dinners from one manufacturer and finds calories with calories, with a boxplot showing no outliers and little skew. A student proposes a two-sample test, treating the 22 dinners as one sample and the 600-calorie guideline as the other. (a) Explain what is wrong with that and name the correct procedure. (b) Carry it out at . (c) Interpret the p-value in context.
Show the worked solution
(a) Why the two-sample idea fails. A two-sample procedure needs two sets of data, each with its own , mean, and standard deviation, so that the sampling variability of both can enter the standard error. The 600-calorie guideline is a fixed published number with no sample behind it, so it contributes no variability. The correct procedure is a one-sample test for a mean, with .
(b) State. Let be the true mean calorie content of this manufacturer's frozen dinners. Because the guideline is an upper limit and the concern is exceeding it, versus , with .
(b) Plan. Random: the 22 dinners are a random sample. 10%: 22 is less than 10% of all dinners this manufacturer produces. Normal/large sample: , so lean on the boxplot, which shows no outliers and little skew.
(b) Compute. calories, so with .
(b) p-value and decision. Right-tailed. In the row, falls between (tail ) and (tail ), so the p-value is between and . Technology gives . That is below , so reject . There is convincing evidence that this manufacturer's dinners average more than 600 calories.
(c) Interpret the p-value. If the true mean calorie content really were exactly 600, only about 2.07% of random samples of 22 dinners would produce a sample mean of 638 calories or higher by chance alone. A result that unlikely under the guideline is what makes the evidence against it convincing. Because was rejected, a wrong decision here would be a Type I error: reporting that the dinners exceed the guideline when in truth they average 600.
(a) A published guideline is a number, not a sample, so it carries no sampling variability; use a one-sample test with . (b) , , , one-sided p-value (table bracket to ). Reject ; convincing evidence the true mean exceeds 600 calories. (c) If the true mean were 600, only about 2.07% of samples of 22 would average 638 calories or more.
Problem 7
Two weather stations report weekly rainfall, in millimeters. From an independent random sample of 12 weeks at Station A and an independent random sample of 15 weeks at Station B, an analyst reports a 95% confidence interval for of , built with the conservative degrees of freedom. (a) Name the procedure and give the degrees of freedom and used. (b) Recover the observed difference in sample means, the margin of error, and the standard error. (c) State what a two-sided test of at would conclude, without computing a p-value. (d) A colleague writes that the study proves the two stations get the same mean weekly rainfall. Correct that.
Show the worked solution
(a) Procedure and df. Two independent random samples of different weeks at different stations, with an interval requested, so this is a two-sample interval for a difference of means. Conservative , and the 95% column of the -table at gives .
(b) Center and margin. The observed difference sits at the midpoint: mm. The margin of error is half the width: mm. So the interval is .
(b) Standard error. Since , invert it: mm.
(c) The test decision. The interval contains 0, so a true difference of zero is a plausible value at 95% confidence. A two-sided test at and a 95% interval always agree, so the test fails to reject : there is not convincing evidence that the two stations differ in true mean weekly rainfall. As a cross-check, at is nowhere near the needed to reject.
(d) Correct the colleague. Failing to reject a null hypothesis is not proof that it is true. The interval runs from to mm, so a 4 mm advantage for Station A is still a plausible value for the true difference, and so is exactly zero. An interval marks the values the data cannot rule out, not values the data support equally: the estimate is mm, which sits standard errors from zero and standard errors from 4 mm, and neither distance reaches the that would push a value outside the interval. The honest statement is that this study did not detect a difference, and with 12 and 15 weeks of data it lacked the precision to detect a modest one. A wider net of weeks would narrow the interval and settle the question.
(a) Two-sample interval, conservative , . (b) mm, mm, mm. (c) The interval contains 0, so the matching two-sided test at fails to reject . (d) Failing to detect a difference is not proof of no difference; a gap as large as 4 mm is still a plausible value.
Problem 8
A teacher wants to know whether reading medium affects comprehension. Two designs are proposed.
Design A. Twenty students are randomly assigned, 10 to read a passage on paper and 10 to read a different but comparable passage on screen. Design B. Ten students each read two comparable passages, one on paper and one on screen, with the order and the passage assignment randomized for each student.
(a) Name the procedure each design calls for. (b) Design B is run, and the differences (paper score minus screen score) are . Test at whether reading medium affects mean comprehension. (c) Explain why Design B is more likely than Design A to detect a real difference of this size. (d) State the scope of inference.
Show the worked solution
(a) Match design to procedure. Design A splits 20 different students into two groups with no link between any pair of them, so it calls for a two-sample test with conservative . Design B has each student read in both media, so the two scores for student 6 belong together: paired test on 10 differences, also with . Same here by coincidence, but very different procedures.
(b) State. Let be the true mean difference in comprehension score (paper minus screen) for students like these. versus , with . Two-sided, because 'affects' names no direction.
(b) Plan. Paired test on the 10 differences. Random: medium order and passage were randomly assigned within each student, which is the form the random condition takes in this experiment. The 10% condition does not apply, since the 10 students are not a sample drawn from a population of pairs. Normal/large sample: , so check the differences: sorted they run , which is single-peaked with no strong skew, and with and the fences at and flag no outliers.
(b) Compute. points. The squared deviations from 2 are , summing to , so . Then and with .
(b) p-value and conclusion. Two-sided with . In the row that clears , the critical value for a tail of , so one tail is below and the two-sided p-value is below . Technology gives . Reject : there is convincing evidence that reading medium affects mean comprehension, with paper scoring higher here. The matching 95% interval, points, excludes 0 and agrees.
(c) Why Design B is more sensitive. Students differ enormously from one another in reading ability, and Design A's standard error is built from all of that between-student variation. Design B measures each student against themselves, so ability cancels out of every difference and the standard error is built only from how a single student's score changes with the medium. Here points, far smaller than the spread of comprehension scores across students would be, which is what makes a 2-point effect detectable with only 10 students.
(d) Scope. Medium was randomly assigned within each student, so a cause-and-effect conclusion about reading medium is justified for these students. But the 10 students were not randomly selected from any larger population, so the finding does not generalize beyond them. Random assignment buys causation; random selection buys generalization, and only the first is present.
(a) Design A: two-sample test, conservative . Design B: paired test on 10 differences, . (b) , , , , , two-sided p-value (below from the table); reject , and the 95% interval agrees. (c) Pairing removes between-student ability differences from the standard error, so a 2-point effect stands out. (d) Causal for these students because of random assignment, but not generalizable, because they were not randomly selected.