Selection bias
By Jude Wallis · Updated
Selection bias is bias created by how the sample was chosen, leaving the individuals studied systematically different from the population.
Selection bias is a property of the sampling method, not of any one sample. It is present when the chance of being included is related to whatever is being measured, so the statistic is centered somewhere other than the parameter no matter how many times the method is run. A description of it is only finished once it names a direction: too high or too low, and for what reason.
Put numbers on one. A school has 600 students, 200 seniors and 400 underclassmen, and 70 percent of seniors against 40 percent of the others want a later start time, so the true proportion is . Surveys handed out at the cafeteria door reach only 25 percent of seniors, who leave campus at lunch, and 90 percent of everyone else. That covers students, of whom want the later start, so the method is aimed at , about 6.3 percentage points low. Every sample it produces is centered on 0.437.
"Our sample came out 60 percent female, so it is biased" is a different claim and usually a wrong one. One lopsided sample is evidence of nothing. In a simple random sample of 100 from a population split evenly, women reach 60 or more about 2.8 percent of the time, and one sex or the other does about 5.7 percent of the time. Bias is diagnosed by asking who the method could never reach, not by inspecting the sample it happened to give.
Selection bias is settled before anyone is contacted. If someone was chosen properly and then failed to answer, that is a later leak with its own name: selection bias vs nonresponse bias.
Topic 1.12 names four biases, voluntary response, undercoverage, nonresponse, and response bias. Selection bias is the umbrella over the ones that act while the sample is still being chosen.
Where this comes up
More collecting data and study design terms, or browse the full statistics glossary.