Population

By Jude Wallis · Published

A population is the entire group of individuals or objects you want to study and draw conclusions about.

The population is fixed by the question, not by the data. Defining one means writing a membership rule precise enough to sort any individual in or out: not "voters" but "the 12,400 people registered to vote in this county on October 1". A number computed from the whole population is a parameter, written μ\mu (mu) for a mean, pp for a proportion, σ\sigma (sigma) for a standard deviation. A parameter is one fixed number, usually unknown, and it does not move when you take a different sample.

Stay with that county. The parameter is pp, the proportion of all 12,400 registered voters who plan to vote yes. A random sample of 600 gives p^=0.52\hat{p} = 0.52 (p-hat), an estimate of pp and not pp itself. The population size enters in one place only: the sample is 600/12400=0.048600/12400 = 0.048 of the population, under the 10 percent ceiling the standard error formula needs, so (0.52)(0.48)/600=0.020\sqrt{(0.52)(0.48)/600} = 0.020 can stand as the standard error.

The wrong sentence is short and common: "the population is the 600 voters who were surveyed." Those 600 are the sample. The population is the group the conclusion is about, and it exists whether or not anyone measures it. Two smaller slips travel with it. A population need not be people; it can be 4,000 laptop batteries, or every 20-minute interval in a factory's day. And a population is not everyone, because it stops exactly where the question stops, so this survey says nothing about the state.

One piece of intuition is worth killing off. A bigger population does not need a bigger sample. Precision comes from nn; apart from that 10 percent check, the population size appears nowhere in p(1p)/n\sqrt{p(1-p)/n}, so a county of 12,400 and a country of 300 million take about the same sample for the same margin of error.

The population you want and the population your method can reach are different things. The list actually drawn from is the sampling frame, and anyone in the population but off that list has no chance of selection. Random sampling is topic 1.11 in Unit 1.

Where this comes up

54 pages on the site use this term.

More collecting data and study design terms, or browse the full statistics glossary.