Nonresponse Bias vs Response Bias
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Nonresponse bias
Collecting data and study design
Nonresponse bias occurs when people chosen for the sample do not respond, and those who do respond differ systematically from those who do not.
Nonresponse bias needs two conditions at once, and students usually check only the first. Selected people fail to give data, and those who responded differ from those who did not in a way that matters for the question. A low response rate by itself is not the bias, only the opening that lets it in: if the silent group would have answered like the responders, the estimate is not pushed anywhere.
A district emails a survey to a simple random sample of 250 teachers asking whether they have enough classroom supplies. 100 reply and 65 say yes, so the reported statistic is (p-hat) on a response rate of . Suppose that among the 150 who never replied, only 45 would have said yes, a rate of 0.30. The truth across the 250 selected teachers is . The survey overstates supply adequacy by 0.21, and nothing in the returned data reveals it.
Here is the sentence that costs points: "only 100 responded, so treat it as a sample of 100." That describes the sample size correctly and the sample wrongly. The 100 are not a random subset of the 250; they are the subset that chose to reply, so a standard error built from describes scatter around 0.65, a center in the wrong place that a wider interval does not move. Emailing 1,000 teachers changes nothing: at the same reply rate the lean returns, measured more precisely.
Nonresponse happens after selection, and that is what separates it from its neighbors. All 250 teachers were on the list and could have answered. Had the list covered only full-time staff, part-timers would have had no chance of selection, which is undercoverage. Had all 250 replied while shading answers toward what an administrator wanted to hear, that would be response bias.
The defense is procedural, not statistical: follow-up contacts, a shorter instrument, and reporting the response rate next to the estimate. Topic 1.12 also wants a direction: the teachers who answered are plausibly those with the strongest feelings about supplies.
Response bias
Collecting data and study design
Response bias occurs when respondents give systematically inaccurate answers, for example due to confusing question wording or pressure to answer a certain way.
Response bias lives in the answers, not in the sample. The right people were reached, they responded, and what they said leans away from the truth in a consistent direction. The usual causes are leading or confusing wording, a sensitive question asked without anonymity, an interviewer whose presence changes what people will say, and faulty recall. Because the lean sits in the instrument, a flawless random sample with a 100 percent response rate can still return a badly wrong number.
A coach asks each of her 200 athletes face to face whether they skipped a scheduled workout last month, and 18 say yes, a reported rate of . The same 200 athletes then answer the same question on an anonymous form, and 62 say yes, or . The gap of 0.22 cannot be sampling variability, because it is the same 200 people both times. It is a property of how the question was asked.
"The athletes lied, so the data are just noisy" packs two errors into one clause. Noise is symmetric and averages out across samples, while under-reporting an embarrassing behavior pushes every repetition of that survey the same way, which is exactly what makes it bias. Blaming the respondents also misplaces the repair. What differed between 0.09 and 0.31 was the format of the question, so that is where the fix belongs: neutral wording, anonymity, or a measurement that does not rely on self-report at all.
The direction is usually predictable. Socially approved behavior gets over-reported and disapproved behavior under-reported, so self-reported exercise and voting run high while self-reported drinking and cheating run low. A leading question pushes answers toward whatever it hints at. Naming which way, in the context of the study, is the half of the answer graders look for.
Two things response bias is not. It is not nonresponse bias, because everyone here answered. It is not undercoverage, because everyone here was reachable. It also survives a perfect sampling design, since fixing who ends up in the sample does nothing to the question those people are handed.