Statistic

By Jude Wallis · Published

A statistic is a numerical value computed from sample data, used to estimate a corresponding population parameter.

A statistic is a number computed from sample data alone. Because its value depends on which individuals were drawn, it changes from one sample to the next, and that movement is what separates it from a parameter, which stays put. Statistics take Roman letters or hats: xˉ\bar{x} (x-bar) for the sample mean, ss for the sample standard deviation, p^\hat{p} (p-hat) for the sample proportion, rr for the sample correlation. Each is paired with the parameter it estimates, μ\mu, σ\sigma, pp, and ρ\rho in turn.

Suppose 60 percent of a large population would vote yes, so p=0.60p = 0.60. Draw a random sample of 80 and find 52 yes votes, giving p^=5280=0.65\hat{p} = \frac{52}{80} = 0.65. Draw a second sample of 80 and find 44, giving p^=4480=0.55\hat{p} = \frac{44}{80} = 0.55. Two different statistics, one unchanged parameter, and no mistake in either sample. That spread is sampling variability, and the sampling distribution of p^\hat{p} describes it.

"The sample proportion is 0.65, so the population proportion is 0.65" is the error to name out loud. A statistic is an estimate and it is almost never exactly right, so the honest version attaches an interval or a standard error to that 0.65. The written form of the same mistake is putting p=0.65p = 0.65 where p^=0.65\hat{p} = 0.65 belongs. That collapse wrecks everything downstream, because a test compares an observed p^\hat{p} against a hypothesized pp, and no comparison is left once both symbols name one quantity.

A statistic has two lives. Before the sample is drawn it is a random variable with a distribution of its own, which is what makes phrases like the mean and standard deviation of p^\hat{p} meaningful. Once the sample is in hand it is a single number.

Two boundaries. A value computed from an entire population is not a statistic, however it was calculated, and a statistic from a badly chosen sample is still a statistic, just a poor estimator. Topic 3.1 grades estimators on exactly those two axes: whether they are centered on the parameter, and how much they vary.

Where this comes up

192 pages on the site use this term.

More collecting data and study design terms, or browse the full statistics glossary.