Effect Size vs Statistical Significance
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Effect size
Hypothesis testing
Effect size measures how big a difference is, in units you can interpret, separately from whether that difference is statistically significant.
An effect size is a measure of magnitude. In its plainest form it is the raw difference the study estimated: 3 points, 0.4 mmHg, 5 percentage points. A standardized version divides that difference by a standard deviation, which strips the units so effects measured on different scales can be lined up against each other. Either form answers how much. A p-value answers a different question and cannot be converted into this one.
Suppose a review program raises a mean score by 0.4 points on a 100-point test, with points. The raw effect is 0.4 points and the standardized effect is standard deviations. Now hold both of those fixed and change only the sample size. At the one-sample statistic is , with a two-sided p-value of 0.424. At the same 0.4 points gives and a p-value of 0.00006. The effect size is 0.04 in both rows.
That is why "the p-value was smaller, so the effect is bigger" cannot be rescued. Those two p-values, 0.424 and 0.00006, describe the identical 0.4-point gain. A p-value blends the size of the effect with the amount of data into a single number, and nothing unblends it, so p-values can never rank two findings by size. Compare the estimates and their intervals instead.
An effect size is a size, and only that. Whether 0.4 points is worth acting on is practical significance, a judgment about context that no formula produces. How firmly the study pinned 0.4 down is the confidence interval. A standardized effect also hides its units, so 0.04 standard deviations means nothing until someone says what a standard deviation of test scores is worth.
No topic in the Fall 2026 topic list is named effect size. The same work is asked for under different wording, in the justify-a-claim topics: 3.11 for a difference between two proportions and 4.8 for a difference between two means hand you a confidence interval for that difference and ask what claim it supports.
Statistical significance
Hypothesis testing
A result is statistically significant when its p-value is at or below alpha: a result at least this extreme would be unlikely if the null hypothesis were true.
A result is statistically significant at level (alpha) when its p-value is at or below , with fixed before the data arrive. The whole claim is about surprise under the null hypothesis: data at least this extreme would be unlikely if were true. It is not a claim about how big the effect is, whether it matters, or how likely is to be true.
Sample size is what makes that a narrow claim rather than an impressive one. Test against and hold the sample result fixed at (p-hat), one percentage point off. At the standard error is , so and the p-value is 0.527, which is nothing at all. At the standard error falls to 0.00158, , and the p-value is under 0.0001. The 51 to 49 split never moved. Only the amount of data did, and the crossing happens at .
So the sentence to distrust is "the result was significant, so the effect is large." Significance measures a gap in standard errors, and the standard error shrinks like , which means a fixed gap reads as stronger and stronger evidence as data pile up. The mirror error is "the result was not significant, so there is no effect." A small study that fails to reject often could not have detected an effect of ordinary size in the first place.
Significance is also a verdict at one chosen threshold, so a p-value of 0.08 is significant at and not at , with nothing about the data changing in between. Report the estimate and a confidence interval next to the verdict: the size of the effect is its effect size, and whether that size is worth acting on is practical significance.