Justifying claims from a mean CI: practice
By Jude Wallis · Published
These 8 problems ask what a confidence interval for a mean actually licenses you to claim. You decide whether a stated value stays plausible, read a paired-difference interval that contains 0, and write the justification for a named audience. Five turn on paired versus two-sample.
AP Statistics: Unit 4 (topics 4.3 Justifying a Claim Based on a Confidence Interval for a Population Mean or Population Mean Difference, 4.8 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means). Topics 4.3 and 4.8 of the Fall 2026 AP Statistics course are where a mean interval stops being arithmetic and becomes an argument: which values stay plausible, what the interval does not establish, and how the conclusion is written in context. Topic 4.3 covers one sample and paired differences, which is why its title names the population mean difference alongside the population mean; topic 4.8 covers two independent samples. The intervals themselves come from topics 4.2 and 4.7.
The one rule these problems run on
A confidence interval for a mean is a list of plausible values for the parameter. That single idea settles every justification question on this material, and it comes in two forms depending on what you are being asked.
- A claimed value. Someone names a number for ("mu"), the population mean: a label says 500 mg, a manufacturer says 0.5 mm. If that number sits inside the interval it stays plausible, so the data give no convincing evidence against the claim. If it sits outside, it is not plausible at that confidence level, so the data do give convincing evidence against it, and which side it falls on tells you the direction.
- A difference. When the parameter is , the true mean difference for paired data, or , the difference between two population means, the claimed value that matters is almost always 0. An interval that excludes 0 is convincing evidence of a real difference and its sign says which way. An interval that contains 0 is not evidence of a difference, and it is also not evidence that the two means are equal.
Problems 1 and 2 work the claimed-value form, problems 3 and 6 work intervals that contain 0, problems 4, 5 and 8 work intervals that exclude it, and problem 7 asks you to write the whole justification for someone who has never taken the course. For the construction side, see confidence interval for means practice and how to interpret a confidence interval for a mean.
Paired or two-sample, decided before anything else
Every difference-of-means problem forks here, and the fork is about the design, not the numbers. Ask one question: does each observation in the first group have a specific partner in the second?
- Paired. The same runner in both shoes, the same water sample run through both instruments, the same route driven under both routing systems, or twins split between treatments. Subtract within each pair, get one column of differences, and run the ordinary one-sample interval on that column with . The parameter is .
- Two independent samples. Different customers at two branches, different plots given two fertilizers, separate random samples from two suppliers. The parameter is , the standard error is , and the exam accepts the conservative or the decimal technology reports.
Getting this wrong is not a rounding issue. Problem 4 runs one paired data set through both procedures: the point estimate is 3.889 seconds either way, but the paired interval is and the two-sample interval is . Same data, opposite conclusions. Pairing removes the runner-to-runner variation from the standard error, and the two-sample procedure does not even apply here, because its independence condition fails the moment the same nine people appear in both columns. Read paired vs two-sample t test and when to use a paired t test if that split is not automatic yet.
Writing the justification so it scores
A justification is a short argument with three named parts: the interval, the value being judged, and a conclusion about the parameter in context. Two templates carry almost all of it.
- Excludes the value. "The 95% confidence interval for the true mean reduction is 0.207 mm to 0.473 mm. Because 0 is not in this interval, there is convincing evidence that the coating reduces mean corrosion depth."
- Contains the value. "The 95% confidence interval for the true mean difference is to 0.232 mg/L. Because 0 is in this interval, a mean difference of zero remains plausible, so there is not convincing evidence that the two methods read differently on average."
Four endings lose credit every time, and problems 1, 3 and 7 are built on them. Writing that the interval proves a value, or that the two means are equal, treats failing to rule something out as having established it. Writing that 95% of individual observations lie between the bounds swaps the mean for the spread of individuals, and the interval is far too narrow for that job, because the standard error is smaller than by a factor of . Attaching a probability to the finished interval ignores that both bounds and are now fixed numbers, so it either captures or it does not; the confidence level describes how often the method succeeds across repeated samples. And a conclusion phrased as "accept" rather than "not convincing evidence" claims more than any interval supports.
Two limits sit outside the arithmetic. An interval says nothing about cause unless treatments were randomly assigned, and it says nothing about a wider population unless the units were randomly selected from one. Problem 5 turns on exactly that, and problems 4, 7 and 8 each earn their causal wording from a randomization inside the design. The four combinations are worked through in can you generalize results.
Problem 1
An orchard weighs a random sample of 30 apples from one harvest and reports a 95% confidence interval for the true mean weight of grams. A dotplot of the 30 weights is roughly symmetric with no outliers.
(a) Find the sample mean and the margin of error.
Say whether each statement below is justified by this interval, and correct the ones that are not.
- About 95% of apples in this harvest weigh between 152.4 g and 161.6 g.
- There is a 95% probability that the true mean weight is between 152.4 g and 161.6 g.
- We are 95% confident that the interval from 152.4 g to 161.6 g captures the true mean weight of all apples in this harvest.
- A true mean weight of 158 g is plausible.
- The interval proves that the true mean weight is not 150 g.
- There is convincing evidence that the true mean weight of apples in this harvest exceeds 150 g.
Show the worked solution
(a) The sample mean sits at the center: ("x-bar") grams. The margin of error is half the width: grams. So the interval is , built at .
Statement 1 is not justified. The interval estimates , the mean weight of all apples in the harvest, not the spread of individual apples. Its half-width 4.6 g is built from the standard error , which is smaller than the apple-to-apple standard deviation by a factor of . A range holding 95% of individual apples would be several times wider.
Statement 2 is not justified. Both bounds are now fixed numbers and is a fixed number, so this interval either captures or it does not. The correct version moves the 95% onto the method: if the orchard repeated this sampling many times with 30 apples each and built an interval from every sample, about 95% of those intervals would capture the true mean weight.
Statement 3 is justified. It names the confidence level, both bounds with units, the parameter, and the population. This is the sentence the exam wants.
Statement 4 is justified. Since , the value 158 g lies inside the interval, so it remains a plausible value for the true mean weight. Note what it does not say: 158 is no more supported than 155 or 160, which are equally inside.
Statement 5 is not justified as written. An interval never proves anything about a single value. Say instead that 150 g lies outside the interval, so it is not a plausible value for the true mean weight at the 95% level.
Statement 6 is justified. Every value in the interval exceeds 150, and 150 sits below the lower bound of 152.4, so there is convincing evidence that the true mean weight is greater than 150 g. Statements 5 and 6 refer to the same fact about 150; only statement 6 words it as the evidence it is.
(a) g and g. Statements 3, 4 and 6 are justified. Statement 1 describes individual apples instead of the mean, statement 2 attaches a probability to a finished interval whose bounds are already fixed, and statement 5 says "proves" when the interval only shows that 150 g is not a plausible value at this confidence level.
Problem 2
A supplement label states that each capsule contains 500 mg of magnesium. A testing lab assays a random sample of 20 capsules from Brand A and finds mg with mg. A dotplot of the 20 assays is roughly symmetric with no outliers.
(a) Construct a 95% confidence interval for the true mean magnesium content of Brand A capsules, and say whether the labeled 500 mg is plausible.
(b) The lab assays 20 capsules of Brand B, which carries the same 500 mg label, and finds mg with the same mg. Build that interval and answer the same question.
(c) A colleague reads part (b) and concludes that the interval shows Brand B capsules really do average 500 mg. Correct him.
Show the worked solution
Plan. One-sample interval for a mean, since is unknown and stands in for it. Random: each brand's 20 capsules are a random sample. 10%: 20 capsules is well under 10% of all capsules the brand produces. Normal or large sample: , so the stated symmetric dotplot with no outliers is what carries the condition. Both brands share and , so they share a standard error and a critical value.
Common arithmetic. mg. With , the 95% column of the -table gives , so mg for both brands.
(a) Brand A: mg. We are 95% confident this interval captures the true mean magnesium content of Brand A capsules.
(a) Justify. The labeled value 500 mg lies above the upper bound 498.148, so it is outside the interval and is not a plausible value for the true mean. There is convincing evidence that Brand A capsules average less than the labeled 500 mg, by somewhere between about 1.9 mg and 12.1 mg.
(b) Brand B: mg. Now , so 500 mg lies inside the interval and remains plausible. There is not convincing evidence against Brand B's label.
(c) Correct him. Failing to rule a value out is not the same as establishing it. The interval leaves every value from 490.852 mg to 501.148 mg plausible, and 494 mg and 499 mg sit in there on exactly the same footing as 500 mg. The sample mean 496 mg is in fact the single best estimate, not 500.
Notice the size of what separated the two answers. The two sample means differ by 3 mg, less than a third of one standard deviation, yet one brand's label is contradicted and the other's is not. A conclusion that flips on a 3 mg shift is a reason to report the interval itself rather than only the verdict.
(a) mg, , , mg, giving mg. 500 mg is above the interval, so it is not plausible and there is convincing evidence Brand A averages under its label. (b) mg contains 500, so the label stays plausible and there is not convincing evidence against it. (c) The interval rules values out, it never establishes one; 500 mg is only one of many plausible values, and 496 mg is the best estimate.
Problem 3
An environmental lab wants to know whether two instruments read nitrate concentration differently. Each of 8 randomly selected water samples is run through both instruments. Concentrations, in mg/L:
| Sample | Instrument A | Instrument B |
|---|---|---|
| 1 | 8.2 | 8.0 |
| 2 | 7.5 | 7.7 |
| 3 | 9.1 | 8.9 |
| 4 | 6.8 | 7.0 |
| 5 | 8.7 | 8.4 |
| 6 | 7.9 | 8.1 |
| 7 | 9.4 | 9.1 |
| 8 | 8.0 | 8.1 |
A dotplot of the differences is roughly symmetric with no outliers.
(a) Explain why this is a paired design, and name the parameter.
(b) Construct a 95% confidence interval for that parameter using .
(c) The lab director writes: "The interval contains 0, so we have shown the two instruments agree." Rewrite the conclusion correctly and explain what is wrong with hers.
Show the worked solution
(a) Each water sample is measured by both instruments, so every reading from A has one specific partner from B: the reading taken on that same sample. That is the definition of pairing, and it is the right design here, because water samples differ a lot from each other and pairing removes that variation from the comparison. The parameter is , the true mean difference in nitrate reading (Instrument A minus Instrument B) for water samples measured by both instruments.
(b) Collapse to one column. gives mg/L. The sum is , so mg/L.
(b) Standard deviation of the differences. Deviations from 0.0375 are , and their squares sum to . So mg/L.
(b) Conditions and interval. Random: the 8 water samples were randomly selected. 10%: 8 is under 10% of the water samples available. Normal or large sample: , so the roughly symmetric dotplot of differences with no outliers carries it. Then mg/L, and with the 95% column gives , so mg/L.
(b) Interval: mg/L. We are 95% confident this interval captures the true mean difference in reading between the two instruments.
(c) What is wrong. The interval containing 0 means a mean difference of zero has not been ruled out; it does not mean zero has been shown. The same interval leaves and mg/L just as plausible as 0, so agreement is one of many surviving possibilities, not a finding. Failing to detect a difference and demonstrating there is none are different claims, and only the first is supported.
(c) Correct version: because 0 lies inside the interval from to mg/L, a true mean difference of zero remains plausible, so these data do not give convincing evidence that the two instruments read differently on average. If the lab needs to establish agreement to within some tolerance, it needs a much narrower interval, which means more paired samples, not a different conclusion from these eight.
(a) Paired, because both instruments measure the same 8 water samples; the parameter is , the true mean difference (A minus B). (b) , , , , , , so the 95% interval is mg/L. (c) The interval contains 0, so a difference of zero stays plausible and there is not convincing evidence the instruments differ. It does not show they agree: values like and mg/L are equally plausible.
Problem 4
Nine randomly selected club runners each run the same course twice, once in their usual shoe and once in a new model, with the order of the two runs randomized. Times, in seconds:
| Runner | Usual | New |
|---|---|---|
| 1 | 248 | 244 |
| 2 | 262 | 259 |
| 3 | 231 | 227 |
| 4 | 275 | 270 |
| 5 | 219 | 216 |
| 6 | 254 | 249 |
| 7 | 240 | 237 |
| 8 | 266 | 262 |
| 9 | 227 | 223 |
A dotplot of the differences is roughly symmetric with no outliers, and so is each column of times.
(a) Construct a 95% confidence interval for the true mean difference using , and justify a claim about the new shoe.
(b) A student ignores the pairing and treats the 18 times as two independent samples. The summary statistics are with for the usual shoe and with for the new one. Build that interval with the conservative and state the conclusion it would produce.
(c) Both intervals are centered at the same number. Explain why they disagree, and say which analysis is valid here.
Show the worked solution
(a) Differences : seconds. The sum is 35, so seconds.
(a) Spread of the differences. Deviations from 3.8889 are , whose squares sum to . So seconds.
(a) Conditions and interval. Random: the 9 runners are a random sample of club runners, and the order of the two runs was randomized within each runner. 10%: 9 is under 10% of the club. Normal or large sample: , so the roughly symmetric dotplot of differences with no outliers carries it. Then s, and at the 95% column gives , so s.
(a) Interval: seconds. Because 0 is not in this interval and the whole interval is positive, there is convincing evidence that runners are faster on this course in the new shoe, by a true mean of somewhere between about 3.3 and 4.5 seconds. Because the order of the two runs was randomized within each runner, the difference is attributable to the shoe rather than to warm-up or fatigue order.
(b) Two-sample version. seconds. The conservative again gives , so s.
(b) Interval: seconds. This interval contains 0, so it would produce the conclusion that there is not convincing evidence of any difference between the two shoes.
(c) Why they disagree. The point estimate is seconds either way, because the mean of the differences equals the difference of the means. Only the standard error changes: 0.2606 s paired against 8.9078 s unpaired, a factor of about 34. Runners differ from one another by roughly 19 seconds on this course, and the two-sample standard error is built from that runner-to-runner spread. Subtracting within each runner cancels it, leaving only the shoe-to-shoe variation, which is well under one second.
(c) Which is valid. The paired analysis. The two-sample procedure requires the two samples to be independent, and they are not: the same nine people supply both columns, so runner 4 being slow shows up in both. The unpaired interval is not merely a wider answer to the same question, it is a procedure whose condition fails on this design. The paired conclusion in part (a) is the one to report.
(a) s, , , , , so the 95% interval is seconds. It excludes 0, so there is convincing evidence the new shoe is faster, by about 3.3 to 4.5 seconds. (b) s and s, giving , which contains 0 and would find no evidence of a difference. (c) Same center, standard errors 34 times apart: pairing removes the roughly 19-second runner-to-runner variation. The two-sample interval is invalid here because its independence condition fails, since the same nine runners appear in both columns.
Problem 5
A software team recruits 36 volunteers and randomly assigns 18 to each of two versions of a scheduling app, then times how long each volunteer takes to complete a fixed set of tasks. Version A gives minutes with ; Version B gives minutes with . Dotplots of both groups are roughly symmetric with no outliers.
(a) Explain why this is a two-sample design and not a paired one.
(b) Construct a 95% confidence interval for using the conservative degrees of freedom, and justify a claim about the two versions.
(c) Can the team conclude that Version B causes faster task completion? Can it conclude that the general public would be faster on Version B?
Show the worked solution
(a) Each volunteer uses exactly one version, so a time from the Version A group has no particular partner in the Version B group. There is nothing to subtract within, and the two groups of 18 are independent by construction of the random assignment. The parameter is , the difference in true mean completion time between Version A and Version B for volunteers like these.
(b) Conditions. Random: volunteers were randomly assigned to the two versions, which is what makes the groups comparable and the two samples independent. Normal or large sample: , so the roughly symmetric dotplots with no outliers carry the condition for both groups. The 10% condition concerns random sampling from a population and is not the relevant check for a randomized experiment on recruited volunteers.
(b) Standard error. minutes.
(b) Critical value and interval. Conservative , so the 95% column gives and minutes. The point estimate is minutes, so the interval is minutes. Technology using reports and , a slightly narrower interval with the same conclusion.
(b) Justify. Because 0 is not in the interval and both bounds are positive, there is convincing evidence that mean completion time is longer on Version A than on Version B, by a true mean of somewhere between about 1.0 and 8.0 minutes. Note how loose that is: the data settle the direction firmly but leave the size uncertain across an eightfold range.
(c) Cause: yes, within this study. Volunteers were randomly assigned to the versions, so the two groups differ systematically only in which version they used, and a difference this large is not plausibly explained by chance assignment alone. That licenses a causal claim about the app version for these volunteers.
(c) Generalization: no. The 36 people were volunteers, not a random sample from any wider population, so nothing here supports extending the result to the general public. Random assignment buys causation; random selection buys generalization, and only the first is present. This is the most commonly missed half of a justification on this material.
(a) Two-sample: each volunteer uses only one version, so no observation has a partner in the other group; the parameter is . (b) min, conservative , , , so the 95% interval is minutes. It excludes 0, so there is convincing evidence Version A takes longer, by about 1.0 to 8.0 minutes on average. (c) Cause yes, generalization no: random assignment supports a causal claim within the study, but volunteers are not a random sample, so the result does not extend to the general public.
Problem 6
A garden center compares tomato seedlings from two seed suppliers. It takes an independent random sample of 12 seedlings from Supplier 1's delivery and an independent random sample of 15 from Supplier 2's, and measures height, in centimeters, after four weeks in the same greenhouse. Supplier 1 gives with ; Supplier 2 gives with . Dotplots of both samples are roughly symmetric with no outliers.
(a) Construct a 95% confidence interval for using the conservative degrees of freedom.
(b) The buyer says the interval shows the two suppliers' seedlings grow to the same mean height. Say what the interval does justify and what it does not.
(c) A colleague suggests the study should have been run as a paired design instead. Was pairing available here?
Show the worked solution
(a) Conditions. Random: two independent random samples, one from each supplier's delivery. 10%: 12 and 15 seedlings are each under 10% of the delivery they came from. Normal or large sample: and are both under 30, so the roughly symmetric dotplots with no outliers carry the condition for both groups.
(a) Standard error. cm.
(a) Critical value and interval. Conservative , so the 95% column gives and cm. The point estimate is cm, so the interval is cm. Technology using gives and , which contains 0 as well.
(b) What it justifies. Zero lies inside the interval, so a true mean difference of zero remains plausible, and there is not convincing evidence that the two suppliers' seedlings differ in mean height after four weeks.
(b) What it does not justify. The buyer's sentence claims the means are the same, which the interval cannot support. Every value from to cm is plausible, so Supplier 1 running a full centimeter shorter and Supplier 2 running nearly two centimeters shorter are both still on the table alongside zero. A difference of zero is one plausible value out of many, and with samples of 12 and 15 the interval is simply too wide to separate them. Report it as no convincing evidence of a difference, never as evidence of no difference.
(c) Pairing was not available. A paired design needs a specific link joining one unit in each group, such as the same subject measured twice or two units matched on a variable before treatment. These are two separate batches of seedlings from two different suppliers, so seedling 3 from Supplier 1 has no partner among Supplier 2's fifteen. Numbering the rows 1 through 12 and 1 through 15 and subtracting would invent a pairing the design does not contain, and it would leave three of Supplier 2's seedlings unmatched.
(c) What would help instead. Pairing is not the lever here; sample size is. The interval is 2.96 cm wide against a point estimate of 0.45 cm, so the study has little chance of detecting a difference this small. Larger samples from both suppliers would shrink the standard error and narrow the interval.
(a) cm, conservative , , , so the 95% interval is cm. (b) It justifies saying there is not convincing evidence of a difference in mean height, since 0 is plausible. It does not justify saying the means are the same: differences anywhere from to cm are equally plausible. (c) No. Pairing needs a specific partner for each unit, and two separate batches from two suppliers have none; larger samples, not pairing, would narrow this interval.
Problem 7
A delivery company tests new routing software on 14 randomly selected routes. Each route is driven once under the old software and once under the new, on comparable days, with the order randomized. For the differences , in minutes, the company reports and . A dotplot of the 14 differences is roughly symmetric with no outliers.
(a) Construct a 95% confidence interval for .
(b) Write the full-credit statistical justification for a claim about the new software.
(c) The operations manager has not taken a statistics course and has to decide whether to buy licenses. Write two or three sentences for her that are honest about what the study shows and how uncertain it is.
(d) Name three sentences that must not appear in either version, and say what is wrong with each.
Show the worked solution
(a) Conditions. Paired design, so this is a one-sample interval on the 14 differences. Random: 14 randomly selected routes, with the order of the two drives randomized within each route. 10%: 14 routes is under 10% of the company's routes. Normal or large sample: , so the roughly symmetric dotplot of differences with no outliers carries it.
(a) Arithmetic. minutes. With , the 95% column of the -table gives , so minutes.
(a) Interval: minutes.
(b) Statistical justification. "We are 95% confident that the interval from 0.967 to 7.433 minutes captures , the true mean reduction in driving time per route (old software minus new). Because 0 is not contained in this interval and both bounds are positive, there is convincing evidence that the new software reduces mean driving time on routes like these. Because the order of the two drives was randomized within each route, the reduction is attributable to the software rather than to which drive came first."
(c) For the manager. "On 14 of our routes, the new software saved an average of 4.2 minutes per route, and the analysis is confident the real average saving is a genuine one rather than luck of the draw. How big that saving is remains loosely pinned down: the data are consistent with anything from about 1 minute to about 7.4 minutes per route, so if the licenses only pay for themselves at 5 minutes a route, this study has not settled the question." That last clause is the part a decision-maker actually needs, and it is invisible if you report only that the result was significant.
(d) "There is a 95% probability that the true saving is between 0.967 and 7.433 minutes." Wrong, because both bounds and are fixed numbers once the data are in, so this interval either captures or it does not. The 95% describes how often intervals built this way capture the parameter across repeated samples.
(d) "The study proves the new software saves 4.2 minutes per route." Wrong twice. An interval does not prove a value, and 4.2 is the sample estimate, not the parameter; the whole point of the interval is that values from about 1 to about 7.4 minutes are all plausible for .
(d) "95% of routes will save between 0.967 and 7.433 minutes." Wrong, because the interval estimates a mean, not the spread of individual routes. Its width comes from , while individual routes vary with minutes, nearly four times as much, so plenty of individual routes will fall outside these bounds and some will get slower.
(a) min, , , , so the 95% interval is minutes. (b) 95% confident the interval from 0.967 to 7.433 minutes captures the true mean reduction per route; 0 is excluded and both bounds are positive, so there is convincing evidence the new software reduces mean driving time, and the randomized drive order attributes it to the software. (c) It saved 4.2 minutes per route on average and the saving looks real, but the data support anything from roughly 1 to 7.4 minutes, so a decision that needs 5 minutes is not settled. (d) Avoid attaching a probability to the finished interval, avoid "proves" and treating 4.2 as the parameter, and avoid describing individual routes, whose spread is rather than .
Problem 8
A coating manufacturer claims its product reduces mean corrosion depth by at least 0.5 mm. A materials lab cuts each of 12 randomly selected steel panels in half, randomly assigns one half of each panel to be coated, exposes all 24 halves to salt spray for 30 days, and measures corrosion depth. For , in millimeters, the lab finds with . A dotplot of the 12 differences is roughly symmetric with no outliers.
(a) Explain why the paired procedure is the right one.
(b) Construct a 95% confidence interval for and say whether there is convincing evidence the coating reduces mean corrosion depth at all.
(c) Is the manufacturer's "at least 0.5 mm" claim plausible at the 95% level?
(d) Redo the interval at 99% confidence. Do your answers to (b) and (c) change?
Show the worked solution
(a) The two halves of a panel come from the same sheet of steel, so they share thickness, alloy, and surface history. That link makes each coated half the natural partner of one uncoated half, and subtracting within a panel removes the panel-to-panel variation from the comparison. The parameter is , the true mean reduction in corrosion depth (uncoated minus coated) for panels treated this way.
(b) Conditions. Random: 12 randomly selected panels, with the coated half chosen at random within each panel. 10%: 12 panels is under 10% of the panels available. Normal or large sample: , so the roughly symmetric dotplot of differences with no outliers carries it.
(b) Arithmetic. mm. With , the 95% column gives , so mm, and the interval is mm.
(b) Justify. Zero lies below the lower bound of 0.207, so 0 is not a plausible value for and every plausible value is a reduction. There is convincing evidence that the coating reduces mean corrosion depth, by a true mean of somewhere between about 0.21 mm and 0.47 mm. Because the coated half was randomly assigned within each panel, the reduction is attributable to the coating.
(c) Now judge the claimed value 0.5, not 0. The claim is that is at least 0.5 mm, and the entire interval from 0.207 to 0.473 lies below 0.5. No value of 0.5 or greater is plausible at this level, so there is convincing evidence against the manufacturer's claim: the coating works, but by less than advertised.
(d) At 99%, only changes. The row gives , so mm and the interval is mm.
(d) Answer to (b) is unchanged. Zero is still below the lower bound of 0.152, so there is still convincing evidence the coating reduces mean corrosion depth.
(d) Answer to (c) flips. The upper bound 0.528 now exceeds 0.5, so and a reduction of 0.5 mm is plausible at the 99% level. There is no longer convincing evidence against the manufacturer's claim.
Take the lesson from (d) seriously. Plausibility is always relative to a confidence level, so a justification that names a value must name the level too. Raising the level widens the interval, which makes the method capture more often across repeated samples but leaves more claimed values standing. Reporting "0.5 mm is not plausible" without "at 95% confidence" hides the fact that the verdict reverses one column over on the -table.
(a) Paired: the two halves of one panel share the steel and its history, so each coated half has a specific partner; the parameter is . (b) mm, , , , so the 95% interval is mm. It excludes 0, so there is convincing evidence the coating reduces mean corrosion depth by about 0.21 to 0.47 mm. (c) No. The whole interval sits below 0.5, so "at least 0.5 mm" is not plausible at 95% and there is convincing evidence against it. (d) At 99%, and the interval is mm. Part (b) is unchanged, but part (c) reverses: 0.5 now lies inside, so the claim is plausible at 99%.