P-value interpretation practice problems
By Jude Wallis · Updated
These 8 problems build p-value interpretation: writing the full contextual sentence, spotting the three standard misreadings, comparing the p-value to alpha, converting one-sided to two-sided, and telling a small p-value apart from a large effect. Five problems ask you to compute the p-value first.
AP Statistics: Unit 3 (topics 3.6 p-Values, 3.7 Carrying Out a Test for a Population Proportion). These problems cover Unit 3 topic 3.6 (p-Values) and the decision rule in topic 3.7 of the Fall 2026 AP Statistics course, where the p-value is defined as a probability computed assuming the null hypothesis is true, and failing to reject is never evidence that the null hypothesis is true.
What these problems build
These 8 problems train one skill: saying exactly what a p-value does and does not tell you. A p-value is the probability, computed by assuming the null hypothesis is true, of getting a test statistic at least as extreme as the one you observed, in the direction of the alternative hypothesis . Two other symbols appear throughout: (read 'p-hat') is the sample proportion, and (alpha) is the significance level you fix before you look at the data.
The set moves from writing a full contextual interpretation, through the three misreadings that show up most often (probability the null is true, probability the result is due to chance, probability you are wrong), to converting between one-sided and two-sided p-values and separating statistical significance from practical importance. Five problems hand you data or a test statistic and make you produce the p-value before you can interpret it.
For the underlying idea, read what does a p-value mean and p-value vs alpha. For picking the cutoff, see how to choose a significance level, and check your tail areas with the p-value calculator. The course statement of the definition is on topic 3.6, p-values, and more sets are on the practice page.
Problem 1
A community solar program states that 45% of its subscribers check the monthly savings dashboard. A staff member takes a random sample of subscribers and runs a one-sample z-test of against , where is the true proportion of all subscribers who check the dashboard. The p-value is . Write a complete interpretation of this p-value in context.
Show the worked solution
Start with the assumption. A p-value is always computed by assuming the null hypothesis is true, so here you assume the true proportion of subscribers who check the dashboard is exactly .
Name the event. The alternative is , so 'at least as extreme' means a sample proportion at least as far above as the one this sample produced.
Attach the number. The p-value is that probability, and the only source of variation it describes is the random sampling.
Assemble the sentence. If the true proportion of subscribers who check the dashboard were , then about 2.71% of random samples of this size would give a sample proportion at least as far above as the one observed.
Check what you did not say. You did not say the probability that is true, and you did not say the probability that the result is 'due to chance'. The probability is about the data, computed under a fixed assumption about .
Assuming the true proportion of subscribers who check the dashboard is , there is a probability that random sampling would produce a sample proportion at least as far above as the one observed.
Problem 2
An urban planner tests whether more than 30% of a city's commuters bike to work at least once a week. She runs a one-sample z-test of against and gets a p-value of . Three students describe it. (a) 'There is a 2.8% chance the null hypothesis is true.' (b) 'There is a 2.8% chance the result happened by chance.' (c) 'If we reject , there is a 2.8% chance we are wrong.' Explain why each is wrong and give the correct statement.
Show the worked solution
(a) The p-value is a probability about data, not about hypotheses. In this framework the true proportion is a fixed unknown number: it either equals or it does not, so it is never assigned a probability. Correct version: assuming , there is a probability of a sample proportion at least this far above .
(b) Every result from a random sample 'happens by chance', so that phrase describes nothing. What the statement drops is the condition that makes the number meaningful, namely that was computed assuming . Correct version: if really were , random sampling alone would produce a result this extreme or more extreme about 2.8% of the time.
(c) The chance that this particular rejection is a mistake is not . A Type I error can only happen when is true, and its long-run rate is the significance level that you chose in advance, not the p-value from one sample.
Spot the shared error. All three sentences delete the phrase 'assuming the null hypothesis is true', which is exactly the phrase that turns a tail area into a p-value.
Test yourself with a rewrite. Any correct interpretation should start with 'if the null hypothesis were true' or 'assuming ', then describe a sample result, then give the probability.
All three are wrong: (a) and (b) either assign probability to a hypothesis or drop the null assumption, and (c) confuses the p-value with . Correct: assuming , about 2.8% of random samples this size would give a sample proportion at least this far above .
Problem 3
A food co-op tests whether the proportion of members who use the bulk bins differs from . The test gives a p-value of . (a) State the decision at . (b) State the decision at . (c) A second co-op runs its own test, sets in advance, and gets a p-value of exactly . What is its decision? (d) Explain what changed between (a) and (b).
Show the worked solution
Recall the rule. Reject when the p-value is less than or equal to ; fail to reject when the p-value is greater than .
(a) Compare. , so reject . There is convincing evidence that the true proportion of members who use the bulk bins differs from .
(b) Compare again. , so fail to reject . At this stricter level there is not convincing evidence that the proportion differs from .
(b) Say what that does not mean. Failing to reject is not evidence that the proportion equals . It means this sample was not surprising enough under to clear a 1% bar.
(c) Apply the rule at the boundary. is true, so the second co-op rejects . The comparison uses 'less than or equal to', so an exact tie goes to rejection.
(d) Identify what moved. Nothing about the data changed between (a) and (b): the sample, the test statistic, and the p-value are identical. Only the cutoff agreed on in advance changed, and a smaller demands stronger evidence before you will reject.
(a) Reject , since . (b) Fail to reject, since . (c) Reject, because the rule is to reject when the p-value is less than or equal to . (d) The evidence is unchanged; only the cutoff changed.
Problem 4
A ferry operator wants evidence that more than 12% of weekday passengers bring a bicycle aboard. From a random sample of weekday passengers, a one-sample z-test gives a test statistic of . (a) Find the p-value. (b) Interpret it in context. (c) What would the p-value be if the alternative had been two-sided?
Show the worked solution
(a) Pick the tail from the alternative. The alternative is 'greater than', so the p-value is the area to the right of under the standard normal curve.
(a) Read the table and subtract. , so the p-value is .
(b) Interpret with the assumption in front. If the true proportion of weekday passengers bringing a bicycle were , about 1.62% of random samples of this size would give a test statistic of or larger.
(c) Use symmetry. The standard normal curve is symmetric about , so a two-sided p-value adds the matching left tail below : p-value .
Compare the two. At both lead to rejecting , The two-sided p-value is exactly twice the one-sided one only when the sample falls on the side the alternative points to, which is what happens here. Problem 5(c) is the other case: when the data go against the alternative, the one-sided p-value climbs above 0.5 and the two-sided p-value is the smaller of the two.
(a) p-value . (b) If the true bicycle rate were , about 1.62% of samples this size would give . (c) Two-sided: .
Problem 5
(a) A two-sided test of gives a p-value of , and the sample proportion came out above . What would the p-value have been for the one-sided alternative ? (b) A right-tailed test gives a p-value of . What is the two-sided p-value for the same data? (c) A researcher tests and gets . Find the p-value and explain what it says.
Show the worked solution
(a) Halve the two-sided value. A two-sided p-value is the sum of two equal tail areas, so one tail is . This works only because the sample fell on the side the alternative points to.
(b) Double the one-sided value. The matching opposite tail has the same area, so the two-sided p-value is .
(c) Choose the tail from the alternative, not from the sign of . The alternative is , so the p-value is the area to the right of .
(c) Compute it. , so .
(c) Read what that means. The sample proportion landed below , the opposite side from where the alternative predicted, so the data provide no support at all for . You fail to reject at any usual .
Keep the warning. For a one-sided test, a p-value above is a signal that the sample went the wrong direction; taking the left tail instead would have given , which is the wrong answer to the question asked.
(a) . (b) . (c) ; the sample fell on the wrong side of for , so there is no evidence for the alternative.
Problem 6
A museum director claims that 35% of visitors come from outside the county. A random sample of 400 visitors contains 132 from outside the county. Test against . (a) Find the test statistic and p-value. (b) State the decision at . (c) A trustee writes, 'the test proves that exactly 35% of visitors come from outside the county.' Correct that sentence.
Show the worked solution
(a) Find the sample proportion. .
(a) Find the standard error using the null value . .
(a) Standardize. .
(a) Find the two-sided p-value. , so the p-value is .
(b) Decide. , so fail to reject . There is not convincing evidence that the true proportion of visitors from outside the county differs from .
(c) Correct the claim. A large p-value says only that a sample proportion of would be unremarkable if were . It gives no evidence that equals , because values such as or would also survive this test.
(c) Restate it properly. 'There is not convincing evidence that the proportion of visitors from outside the county differs from 35%.' A hypothesis test can reject or fail to reject , but it can never establish that is true.
(a) , , two-sided p-value . (b) Fail to reject at . (c) A large p-value never proves the null; it says only that would not be surprising if the true proportion were .
Problem 7
A podcast platform moves the credits to the end of each episode, then checks whether the proportion of plays that reach the end differs from . In a random sample of 40,000 plays, 20,280 reach the end. (a) Find the test statistic and the p-value. (b) State the decision at . (c) A 95% confidence interval for the true completion rate is . Use it to explain why a very small p-value does not mean the change matters.
Show the worked solution
(a) Find the sample proportion. .
(a) Find the standard error using . .
(a) Standardize. .
(a) Find the two-sided p-value. , so one tail is and the p-value is . Technology reports .
(b) Decide. , so reject . There is convincing evidence that the true completion rate differs from .
(c) Size the effect. The estimate is , and the interval says plausible values run from about to . Even the top of that range sits only above one half, roughly one extra completed play per 100.
(c) Explain the small p-value. With the standard error is only , so a gap of is standard errors wide even though it is tiny in real terms. Large samples turn small differences into small p-values.
(c) Separate the two questions. The p-value answers 'could random variation alone plausibly produce a gap this big?' It does not answer 'is the gap large enough to act on?' For that you read the estimate and the interval.
(a) , , two-sided p-value . (b) Reject at . (c) The interval puts the true rate at most about 1.2 percentage points above , so the result is statistically significant but practically trivial.
Problem 8
A regional park installs new trailhead signs asking hikers to pack out their trash. Rangers then take a random sample of 250 trail users and find 212 who carry out all of their trash. Before the signs, the park's rate was . (a) State the hypotheses and check the conditions. (b) Find the test statistic and the p-value. (c) Interpret the p-value in context. (d) State a decision at , then correct this student sentence: 'The p-value of is the probability that the signs had no effect.' (e) Give the decision at and say what that does and does not change.
Show the worked solution
(a) State. Let be the true proportion of trail users at this park who carry out all of their trash after the signs went up. versus .
(a) Check conditions. Random: the 250 trail users are a random sample. 10%: the park logs roughly 12,000 trail users a season, and 250 is well under 10% of that. Large Counts under : and , both at least 10.
(b) Find the sample proportion and standard error. , and .
(b) Standardize and find the tail. . The test is right-tailed, so the p-value is .
(c) Interpret. If the true carry-out rate were still , about 2.87% of random samples of 250 trail users would give a sample proportion at least as far above as .
(d) Decide. , so reject . There is convincing evidence that the true carry-out rate is now greater than .
(d) Correct the sentence. It is wrong twice. The p-value is not a probability about the signs or about being true, and it is computed by assuming the rate is still . Correct version: assuming the rate is , there is a probability of a sample result at least this far above .
(d) Add the design caution. Rejecting says the rate is above ; with no control group, it does not by itself show that the signs caused the increase, since weather or the season could differ from the earlier period.
(e) Decide at the stricter level. , so you fail to reject . What changes is the decision. What does not change is the sample, the test statistic , the p-value , or the strength of the evidence, and failing to reject still does not show the rate equals .
(a) , ; conditions met with and . (b) , , p-value . (c) If the true carry-out rate were still 0.80, about 2.87% of random samples of 250 trail users would give a sample proportion at least as far above 0.80 as 0.848 is. (d) Reject at ; the p-value is the probability of a result this extreme assuming the rate is , not the probability the signs did nothing. (e) At you fail to reject, and only the cutoff changed.