Sampling distribution

By Jude Wallis · Published

A sampling distribution is the distribution of a statistic across all possible samples of the same size drawn from a population.

Three different distributions get called "the distribution" and only one of them is this one. The population distribution holds a value for every individual. The distribution of a sample holds the values in the one sample you actually collected. A sampling distribution is neither: its individuals are entire samples of a fixed size nn, and the number recorded for each is a statistic computed from it, such as xˉ\bar{x} (x-bar) or p^\hat{p} (p-hat).

Small cases can be written out in full. Take the population 2, 4, 6, 8, with μ=5\mu = 5 (mu) and σ=2.2361\sigma = 2.2361 (sigma), and draw samples of size 2 with replacement. There are exactly 16 such samples, and their means come out as 2 once, 3 twice, 4 three times, 5 four times, 6 three times, 7 twice, and 8 once. That list is the sampling distribution of xˉ\bar{x}. Its mean is 5, matching μ\mu, and its standard deviation is 1.5811, which is σ/2\sigma/\sqrt{2}. The population is flat; the sampling distribution is already triangular at n=2n = 2.

The misreading is blunt: "the sampling distribution is the distribution of my sample." A histogram of the 40 numbers you collected is a picture of the sample. As nn grows it comes to look more like the population, skew and all, and it does not narrow. The sampling distribution is the thing that narrows, and σ/n\sigma/\sqrt{n} describes it rather than your data.

Outside enumerable toy cases you never actually build one. Theory stands in for it, the central limit theorem for means and the binomial for counts, or a simulation approximates it: 10,000 simulated sample means draw a close picture of the sampling distribution without being it.

Topic 2.12 introduces sampling distributions. Every confidence interval and every p-value later in the course is a statement read off one.

One sampling distribution gets reported on the news every month. The US unemployment rate is not a census, it is an estimate from a household survey of about sixty thousand homes, which is why it is revised and why it carries a margin of error at all: the unemployment rate.

Where this comes up

35 pages on the site use this term.

More sampling distributions terms, or browse the full statistics glossary.