Statistical significance
By Jude Wallis · Updated
A result is statistically significant when its p-value is at or below alpha: a result at least this extreme would be unlikely if the null hypothesis were true.
A result is statistically significant at level (alpha) when its p-value is at or below , with fixed before the data arrive. The whole claim is about surprise under the null hypothesis: data at least this extreme would be unlikely if were true. It is not a claim about how big the effect is, whether it matters, or how likely is to be true.
Sample size is what makes that a narrow claim rather than an impressive one. Test against and hold the sample result fixed at (p-hat), one percentage point off. At the standard error is , so and the p-value is 0.527, which is nothing at all. At the standard error falls to 0.00158, , and the p-value is under 0.0001. The 51 to 49 split never moved. Only the amount of data did, and the crossing happens at .
So the sentence to distrust is "the result was significant, so the effect is large." Significance measures a gap in standard errors, and the standard error shrinks like , which means a fixed gap reads as stronger and stronger evidence as data pile up. The mirror error is "the result was not significant, so there is no effect." A small study that fails to reject often could not have detected an effect of ordinary size in the first place.
Significance is also a verdict at one chosen threshold, so a p-value of 0.08 is significant at and not at , with nothing about the data changing in between. Report the estimate and a confidence interval next to the verdict: the size of the effect is its effect size, and whether that size is worth acting on is practical significance.
Where this comes up
More hypothesis testing terms, or browse the full statistics glossary.