What Does a P-Value Mean? (Plain-English Guide)

By Jude Wallis · Published

A p-value is the probability of getting a result at least as extreme as the one you observed, assuming the null hypothesis is true. A small p-value means your data would be unlikely under the null hypothesis, so it gives evidence against it.

AP Statistics: Unit 3 (topics 3.6 p-Values, 3.7 Carrying Out a Test for a Population Proportion). This maps to AP Statistics Unit 3 (Inference for Categorical Data: Proportions) in the Fall 2026 course, where topic 3.6 p-Values introduces the p-value and topic 3.7 carries out the one-proportion z-test.

What does a p-value mean?

A p-value is the probability, assuming the null hypothesis is true, of getting a sample result at least as extreme as the one you observed.

Break that into three parts. First, it is a probability, so it lives between 0 and 1. Second, it is computed assuming the null hypothesis is true, which is the starting claim you are testing. Third, "at least as extreme" means your observed result plus every outcome even further from what the null hypothesis predicts.

A small p-value means your data would be surprising if the null hypothesis were true, so it counts as evidence against that hypothesis. A large p-value means your data sit comfortably with the null hypothesis, so you have no strong reason to doubt it.

The null hypothesis is usually written H0H_0 (read "H-naught") and the alternative hypothesis is written HaH_a. For how to state these, see null vs alternative hypothesis.

How to interpret a p-value

To interpret a p-value, state it as a conditional probability in context. Suppose you test whether a coin is fair and get a p-value of 0.03. The correct reading is this: if the coin were truly fair, there would be a 3% chance of seeing a result at least as far from an even split as the one you got.

Notice what the p-value measures. It measures how well your data agree with the null hypothesis, not how likely the null hypothesis itself is. The hypothesis is treated as fixed, and the data are what vary from sample to sample.

You then compare the p-value to a significance level, written α\alpha (the Greek letter alpha), that you pick before collecting data. If the p-value is less than or equal to α\alpha, you reject H0H_0. If it is greater than α\alpha, you fail to reject H0H_0. Failing to reject is not the same as proving H0H_0 true. It only means your data did not give enough evidence against it.

What a p-value is not

Three misreadings show up constantly. Each is wrong for the same reason: the p-value is a probability about the data, not about the hypotheses.

It is not the probability that the null hypothesis is true. The p-value already assumes H0H_0 is true and then asks about the data, so it cannot also be the probability that H0H_0 is true. That would reverse the conditional. What you actually computed is closer to P(data this extremeH0 true)P(\text{data this extreme} \mid H_0 \text{ true}), which is a different quantity from P(H0 truedata)P(H_0 \text{ true} \mid \text{data}).

It is not the probability that your result happened by chance. Every sample result involves chance. A p-value of 0.03 does not mean there is a 3% chance your finding is a fluke and a 97% chance it is real. It only describes how often chance alone, under the null hypothesis, would produce data as extreme as yours.

It is not the probability that you are wrong. A p-value of 0.04 does not mean there is a 4% chance your conclusion is a mistake. The long-run rate of falsely rejecting a true null hypothesis is α\alpha, which you set in advance, not the p-value you happen to observe. See type 1 vs type 2 errors for how those error rates work.

Is a p-value of 0.05 significant?

A p-value of 0.05 is treated as significant under the most common convention, because 0.05 is the significance level people use by default. But 0.05 is a convention, not a law of nature. Nothing physical changes between a p-value of 0.049 and one of 0.051. Only the reject-or-not label flips.

The threshold α\alpha is a choice about how much risk of a false alarm you are willing to accept. A stricter level such as 0.01 makes you demand stronger evidence before rejecting H0H_0, while a looser level such as 0.10 rejects more readily. You should pick α\alpha before you see the data so the decision rule is not shaped by the result.

Because the exact number carries information, report the actual p-value rather than only writing "p < 0.05". A reader can then judge the strength of the evidence instead of seeing a single yes-or-no verdict.

What a p-value less than 0.05 means

A p-value less than 0.05 means that, if the null hypothesis were true, data at least as extreme as yours would occur less than 5% of the time. Under the 0.05 convention that is unusual enough to reject H0H_0 and call the result statistically significant.

Statistically significant is not the same as large or important. With a very big sample, a tiny and unimportant difference can still produce a p-value below 0.05. With a small sample, a real and sizable effect can produce a p-value above 0.05. The p-value reflects both the size of the effect and the amount of data, so read it alongside the actual estimate rather than on its own.

One-tailed vs two-tailed p-values

The tail you use depends on your alternative hypothesis. If HaH_a says the parameter is simply different from the null value (for example Ha:p0.5H_a: p \neq 0.5), the test is two-tailed, and the p-value adds the probability in both tails of the distribution.

If HaH_a points in one direction (for example Ha:p>0.5H_a: p > 0.5), the test is one-tailed, and the p-value uses only the single tail in that direction. For a symmetric distribution such as the standard normal, the one-tailed p-value is half the two-tailed p-value when the observed result falls in the direction the alternative predicts. If the result comes out on the opposite side, the one-tailed p-value is instead larger than 0.5, so the halving rule does not apply.

Choose the direction from the question before you collect data, not after you see which way the result came out. Switching to a one-tailed test because it produces a smaller p-value is a misuse of the method.

How to compute a p-value with a one-proportion z-test

For a claim about a single population proportion, AP Statistics uses the one-proportion z-test (Unit 3, topics 3.6 and 3.7). The steps are:

  1. Write H0:p=p0H_0: p = p_0 and an alternative, where p0p_0 is the hypothesized proportion and pp is the true population proportion.
  2. Check that the sample is random, that np010n p_0 \geq 10 and n(1p0)10n(1 - p_0) \geq 10, and that the sample is less than 10% of the population.
  3. Find the sample proportion p^\hat{p} (read "p-hat"), the fraction of successes in your sample.
  4. Compute the standardized test statistic $z=p^p0p0(1p0)n.z = \frac{\hat{p} - p_0}{\sqrt{\frac{p_0(1 - p_0)}{n}}}.Thedenominatoristhestandarddeviationofthesamplingdistributionof The denominator is the standard deviation of the sampling distribution of \hat{p}when when H_0$ is true.
  5. Turn zz into a p-value using the standard normal distribution and your alternative hypothesis, then compare it to α\alpha.

The worked example below runs these steps with numbers. You can check your arithmetic with the p-value calculator or the proportion z-test calculator.

Computing a two-tailed p-value from a one-proportion z-test

A state official claims that 50% of registered voters support a ballot measure. A random sample of 100 voters finds 60 who support it. Using α=0.05\alpha = 0.05, test whether the true proportion of all registered voters who support the measure differs from 0.50.

  1. State the hypotheses. Let pp be the true proportion of all registered voters who support the measure. H0:p=0.50H_0: p = 0.50 and Ha:p0.50H_a: p \neq 0.50. The significance level is α=0.05\alpha = 0.05.

  2. Check conditions. The sample is random. Using p0=0.50p_0 = 0.50: np0=100(0.50)=5010n p_0 = 100(0.50) = 50 \geq 10 and n(1p0)=100(0.50)=5010n(1 - p_0) = 100(0.50) = 50 \geq 10. The sample of 100 is less than 10% of all registered voters, so the observations are effectively independent, and the Large Counts check above makes the normal approximation reasonable.

  3. Find the sample proportion. p^=60100=0.60\hat{p} = \frac{60}{100} = 0.60.

  4. Find the standard deviation of the sampling distribution under H0H_0. p0(1p0)n=0.50(0.50)100=0.25100=0.0025=0.05\sqrt{\frac{p_0(1 - p_0)}{n}} = \sqrt{\frac{0.50(0.50)}{100}} = \sqrt{\frac{0.25}{100}} = \sqrt{0.0025} = 0.05.

  5. Compute the standardized test statistic. z=0.600.500.05=0.100.05=2.00z = \frac{0.60 - 0.50}{0.05} = \frac{0.10}{0.05} = 2.00.

  6. Find the p-value. The test is two-tailed, so the p-value is P(Z2.00)+P(Z2.00)=2×P(Z2.00)P(Z \leq -2.00) + P(Z \geq 2.00) = 2 \times P(Z \geq 2.00), where ZZ is a standard normal random variable. From the z-table, P(Z<2.00)=0.9772P(Z < 2.00) = 0.9772, so P(Z2.00)=10.9772=0.0228P(Z \geq 2.00) = 1 - 0.9772 = 0.0228. The p-value is 2(0.0228)=0.04562(0.0228) = 0.0456.

  7. Compare and conclude. Because 0.0456<0.050.0456 < 0.05, reject H0H_0. There is convincing statistical evidence that the proportion of voters who support the measure differs from 0.50.

  8. Interpret the p-value. If exactly 50% of voters supported the measure, there would be about a 4.56% chance of getting a sample proportion at least as far from 0.50 as 0.60 is, in either direction.

z=2.00z = 2.00 and the two-tailed p-value is 2(0.0228)=0.04562(0.0228) = 0.0456. Since 0.0456<0.050.0456 < 0.05, reject H0H_0. (If the alternative had instead been Ha:p>0.50H_a: p > 0.50, the one-tailed p-value would be P(Z2.00)=0.0228P(Z \geq 2.00) = 0.0228.)

Frequently asked questions

Does a large p-value prove the null hypothesis is true?

No. A large p-value means your data are consistent with the null hypothesis, not that the null hypothesis is correct. You fail to reject H0H_0, which only signals a lack of evidence against it. A different or larger sample could still turn up evidence later.

Is a smaller p-value always better?

Not by itself. A smaller p-value gives stronger evidence against the null hypothesis, but it says nothing about how big or important the effect is. A large sample can make a trivial difference produce a very small p-value, so always read the p-value next to the actual estimate.

Can a p-value be greater than 1?

No. A p-value is a probability, so it always falls between 0 and 1. If a calculation gives a value above 1, you have added tail areas incorrectly or used the wrong tail for your alternative hypothesis.