How to tell which type of bias a survey has
By Jude Wallis · Published
Ask where the lean entered. If people put themselves in the sample, it is voluntary response bias. If part of the population was never on the list, undercoverage. If selected people never answered, nonresponse. If the answers lean away from the truth, response bias.
AP Statistics: Unit 1 (topics 1.12 Potential Problems with Sampling). Topic 1.12 (Potential Problems with Sampling) is where the Fall 2026 course names voluntary response, undercoverage, nonresponse, and response bias, and asks you to describe how each affects a result; it sits in Unit 1, which is 20% to 30% of the multiple-choice section.
Ask where the lean entered
Every survey moves through four stages, and each named bias enters at exactly one of them. The population is who you want to describe. The list is who your method could actually reach. The sample is who got selected. The data is what those people said.
- Undercoverage bias enters at the list: part of the population was left off it, or made less likely to be picked.
- Voluntary response bias enters at selection: the people in the sample chose themselves by answering an open invitation.
- Nonresponse bias enters after selection: the right people were chosen, and some of them never answered.
- Response bias enters at the data: the right people answered, but their answers lean away from the truth.
In all four cases bias means the same thing. It is a systematic error built into the method that pushes a statistic consistently above or consistently below the parameter it estimates, so the sampling distribution of a sample proportion is centered away from the population proportion . Single samples still bounce around by chance, and some of them land on the far side of , but the center of that bouncing sits on the wrong side and repeating the method never moves it back. Bias belongs to the method, not to one unlucky sample. That is why a bigger sample never fixes it: you get a more precise version of the same wrong answer.
The decision flow: four questions in order
Read the scenario once, then walk these four questions in order. The first yes names the bias you lead with, but keep going after it, because a scenario can trip more than one and the exam wants each one named.
- Did the people in the sample put themselves there? An invitation to call in, click a link, mail back a form, or post a review means the sample is made of volunteers. That is voluntary response bias.
- Was some part of the population unable to be selected? If the frame the researcher drew from left people out, or gave them a smaller chance, that is undercoverage.
- Were people selected properly and then never heard from? Refusals, no one home, blank questions. That is nonresponse.
- Did the chosen people answer, but with answers you would not trust? Leading wording, an interviewer standing there, an embarrassing question. That is response bias.
Work the questions in that order because self-selection is the loudest signal in the wording and settles a scenario fastest, while response bias is the one you can only see once you read the question the subjects were asked. The four stages above tell you where each bias enters; this list tells you which to check first. When two of them apply, name both and say where each one entered. Here is what each one looks like in the wording of a problem.
| Bias | The tell in the scenario | Where it enters | Which way it usually leans |
|---|---|---|---|
| Voluntary response | "listeners were invited to call in", "click here to vote" | who ends up in the sample | toward whoever feels strongly, often the unhappy side |
| Undercoverage | "the list included only", "households with a landline" | who could be selected | toward whatever the covered group thinks |
| Nonresponse | "only 34 of the 100 replied", "no one answered at 12 homes" | who answers | toward whoever bothers to respond |
| Response bias | "the question asked", "asked face to face whether they had ever cheated" | what the answers say | toward the socially acceptable answer, or the way the question points |
Undercoverage vs nonresponse: the pair that costs points
This is the confusion that shows up most, and one question settles it: could that person have been selected at all?
If the answer is no, it is undercoverage. Those people were never in the running. A phone poll drawn from landline numbers cannot reach cell-only households. An online panel cannot reach people with no internet. A roster built on Monday misses students who enrolled on Tuesday.
If the answer is yes, and the person was selected but never gave you data, it is nonresponse. They refused, they were not home, they ignored the email, they skipped that one item on the form.
Undercoverage happens before anyone is contacted. Nonresponse happens after. That is the whole distinction.
The trap is that both leave you with missing people, so "only 34 of the 100 selected students replied" can read like a coverage problem. It is not: all 100 were selected, so the 66 silent students are nonresponse. And notice the second half of the definition. A low response rate by itself is not automatically bias. It becomes nonresponse bias when the people who did not reply differ from the people who did in a way that matters for the question being asked.
Voluntary response vs convenience: who did the choosing
Both are nonrandom ways to get a sample, and the difference is who made the choice.
In a voluntary response sample the researcher opens a door and waits. Radio call-in polls, online opinion polls, comment cards, product reviews. People select themselves, and people with strong feelings select themselves far more often than people who are indifferent.
In a convenience sample the researcher picks whoever is easiest to reach. The first 30 students in the cafeteria, shoppers at one mall entrance, the teacher's own class. Nobody volunteered; the researcher chose them for convenience.
On the exam, convenience sampling is usually named as the method, and the bias you then describe works like undercoverage: every student who was not in that cafeteria had no chance of being selected. So name the method, then say who was shut out and how those people differ. Both methods fail for the same underlying reason, which is that chance did not decide who got in. That is exactly what random selection is for, and the four random methods are compared in simple random vs stratified sampling.
Name the direction, in context, or you lose the point
Naming the type is half the answer. The course asks you to describe how the bias affects the result, and graders want a direction attached to the actual variable in the study. Use this shape:
Because [the affected group] tends to [differ in this specific way], this method will [overstate or understate] [the parameter, in context].
How to find the direction for each type:
- Voluntary response. People who answer an open call feel strongly, so the estimate overstates how common that strong view is. A restaurant's online review score is usually pulled by diners who loved it or hated it, not by the quiet middle.
- Undercoverage. The direction comes from how the missing group differs. If the excluded people are less likely to have the trait, the estimate comes out too high; if they are more likely, too low.
- Nonresponse. Same logic, one stage later. Ask how non-responders differ from responders, then push the estimate toward the responders.
- Response bias. Embarrassing behavior gets under-reported and admirable behavior gets over-reported, so self-reported cheating runs low and self-reported exercise runs high. A leading question pushes answers toward whatever it hints at.
Two sentences that earn nothing: "the results will be biased" and "the sample was too small." The first has no direction and no context; the second is not bias at all. If you truly cannot pin the direction, still say who is missing and how the estimate would move if that group differs, which is the reasoning graders are looking for.
What is not bias
Three things get called bias in student answers and are not.
- Sampling variability. A perfectly good random method gives a different statistic every time you run it. That spread is chance, not a lean. Bias is the method missing the parameter in the same direction over and over.
- A small sample. Small samples are more variable, not more biased. Raising narrows the spread of the sampling distribution and does nothing at all to a systematic lean, which is worked through in does a bigger sample fix bias.
- One odd-looking sample. You cannot diagnose bias from a single result. You diagnose it from the procedure that produced it, which is why exam questions hand you the procedure.
One more boundary. This topic is about surveys and sampling. Confounding, control groups, and placebo effects belong to experimental design, so if the scenario assigns treatments rather than asking questions, you are in a different topic. See experiments vs observational studies for that line, and parameter vs statistic for the vocabulary that makes "the statistic misses the parameter" precise.
Name the bias in six scenarios
Identify the problem with each sampling method and say which way it pushes the result. (a) A magazine prints a questionnaire and asks readers to mail it back; 1,200 do. (b) A city polls households about support for a new transit line by dialing numbers from a landline directory, so cell-only households are not on the list. (c) A principal takes a simple random sample of 100 students and emails a survey; 34 reply. (d) A group asks, "Do you support the wasteful new stadium tax?" (e) An interviewer asks students face to face whether they have ever cheated on a test. (f) A reporter interviews the first 25 people leaving a gym about how much they exercise.
(a) Readers decided for themselves whether to mail the form back, so the sample is made of volunteers. Voluntary response bias. Readers with strong opinions about the magazine's topic are the ones who write in, so the result overstates how common those strong opinions are among readers.
(b) Cell-only households were never on the list, so they had no chance of being selected. Undercoverage. Cell-only households skew younger, and younger residents back transit projects at higher rates than older ones, so the poll reflects only landline-owning residents and understates the proportion of households that support the transit line.
(c) All 100 students were properly selected by an SRS, and 66 simply did not reply. Nonresponse. If the 34 who replied are the students most engaged with school, the survey overstates engagement for the whole student body.
(d) The word "wasteful" tells respondents what to think before they answer. Response bias from leading wording, which pushes the reported support below the true level of support.
(e) Cheating is embarrassing, and the interviewer is standing there. Response bias from a sensitive question asked without anonymity, which pushes the reported cheating rate below the true rate.
(f) The reporter chose the easiest people to reach, so this is a convenience sample, not a random one. Everyone not at the gym had no chance of being interviewed, so the bias works like undercoverage: people leaving a gym exercise far more than the town as a whole, and the estimate of average exercise comes out too high.
(a) voluntary response bias, overstating strong opinions; (b) undercoverage of cell-only households, understating support for the transit line; (c) nonresponse bias, overstating engagement if repliers are the engaged students; (d) response bias from leading wording, understating support; (e) response bias from a sensitive question, understating cheating; (f) convenience sampling, which shuts out non-gym-goers and overstates exercise.
Two biases stacked, with the direction measured
A town has 10,000 registered voters. The clerk has an email address for 6,000 of them and none for the other 4,000. A pollster emails everyone on the list and asks whether they support a ballot measure. Unknown to the pollster, 45% of the 6,000 listed voters support the measure and 20% of the 4,000 unlisted voters support it. (a) Find the true population proportion of supporters. (b) If every listed voter replied, what would the poll report, and which bias is that? (c) Now suppose only 1,500 reply: 900 of the supporters on the list and 600 of the non-supporters. What does the poll report now, and which bias was added?
Count supporters in each group. Listed: supporters. Unlisted: supporters.
(a) Total supporters , so the parameter is .
(b) With every listed voter replying, the poll reports . The 4,000 unlisted voters had no chance of being selected, so this is undercoverage. It overstates support by , because the uncovered group supports the measure at a much lower rate (20% against 45%).
(c) The list holds 2,700 supporters and non-supporters. Responders: 900 supporters and 600 non-supporters, for replies.
Compare response rates: , so about 33% of supporters replied, against , about 18% of non-supporters. Supporters answered at nearly twice the rate, so they are over-represented among the replies. That is nonresponse bias.
The reported statistic is now , which sits above the true proportion. Collecting more replies at these same rates of about 33% and about 18% would not move either lean, but following up with the listed voters who stayed silent would shrink the second one, because it brings in the people who are currently missing: if all 6,000 eventually replied the poll would report 0.45 again. Nothing sent to this list can touch the first lean, since the 4,000 voters with no address are not on it.
(a) . (b) The poll reports 0.45, an undercoverage bias that overstates support by 0.10 because unlisted voters support the measure far less. (c) The poll reports 0.60, adding nonresponse bias on top, since supporters replied at about 33% against about 18% for non-supporters, for a total overstatement of 0.25.
Frequently asked questions
What is the difference between undercoverage and nonresponse bias?
Ask whether the person could have been selected. Undercoverage means no: they were left off the list the sample was drawn from, like cell-only households in a landline poll. Nonresponse means yes: they were selected and then did not answer. Undercoverage happens before contact, nonresponse after.
Is a voluntary response sample the same as a convenience sample?
No, and the difference is who does the choosing. In a voluntary response sample people select themselves by answering an open invitation, such as an online poll. In a convenience sample the researcher picks whoever is easiest to reach. Both are nonrandom, so both are biased, but only voluntary response involves self-selection.
Can one survey have more than one type of bias?
Yes, and real surveys usually do. An email poll can leave out people with no email address, which is undercoverage, then suffer nonresponse from the people it does reach, and use a leading question on top of that. Name each one, say where it enters, and give the direction it pushes the estimate.
Does a larger sample reduce bias?
No. Bias comes from the method, so a larger sample gives a more precise version of the same wrong answer. A bigger sample narrows sampling variability, which is a different problem. Only changing the method, usually by selecting with random chance, reduces bias.