Effect size

By Jude Wallis · Updated

Effect size measures how big a difference is, in units you can interpret, separately from whether that difference is statistically significant.

An effect size is a measure of magnitude. In its plainest form it is the raw difference the study estimated: 3 points, 0.4 mmHg, 5 percentage points. A standardized version divides that difference by a standard deviation, which strips the units so effects measured on different scales can be lined up against each other. Either form answers how much. A p-value answers a different question and cannot be converted into this one.

Suppose a review program raises a mean score by 0.4 points on a 100-point test, with s=10s = 10 points. The raw effect is 0.4 points and the standardized effect is 0.4/10=0.040.4/10 = 0.04 standard deviations. Now hold both of those fixed and change only the sample size. At n=400n = 400 the one-sample tt statistic is 0.4/(10/400)=0.800.4/(10/\sqrt{400}) = 0.80, with a two-sided p-value of 0.424. At n=10,000n = 10{,}000 the same 0.4 points gives t=4.00t = 4.00 and a p-value of 0.00006. The effect size is 0.04 in both rows.

That is why "the p-value was smaller, so the effect is bigger" cannot be rescued. Those two p-values, 0.424 and 0.00006, describe the identical 0.4-point gain. A p-value blends the size of the effect with the amount of data into a single number, and nothing unblends it, so p-values can never rank two findings by size. Compare the estimates and their intervals instead.

An effect size is a size, and only that. Whether 0.4 points is worth acting on is practical significance, a judgment about context that no formula produces. How firmly the study pinned 0.4 down is the confidence interval. A standardized effect also hides its units, so 0.04 standard deviations means nothing until someone says what a standard deviation of test scores is worth.

No topic in the Fall 2026 topic list is named effect size. The same work is asked for under different wording, in the justify-a-claim topics: 3.11 for a difference between two proportions and 4.8 for a difference between two means hand you a confidence interval for that difference and ask what claim it supports.

Where this comes up

More hypothesis testing terms, or browse the full statistics glossary.