Voluntary Response Sample vs Nonresponse Bias

Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.

Voluntary response sample

Collecting data and study design

A voluntary response sample is made up of people who choose to respond on their own, such as to an open online poll or a call-in survey.

In a voluntary response sample nobody gets selected. The researcher opens a channel, a link, a call-in line, a mail-back form, a review page, and the sample is whoever decides to use it. That self-selection is systematic rather than random: people who take the trouble to answer an unsolicited invitation tend to hold stronger views than the people who ignore it, so the sample overrepresents both ends of the opinion scale and thins out the indifferent middle.

A school newspaper posts a link asking the school's 1,600 students whether the dress code should change. 300 answer and 240 of them say yes, so the poll reports p^=240300=0.80\hat{p} = \frac{240}{300} = 0.80 (p-hat). Suppose that among the 1,300 students who ignored the link, 455 would have said yes, a rate of 0.35. The true proportion is 240+4551600=0.434\frac{240 + 455}{1600} = 0.434. The poll is off by more than 0.36 while displaying 300 responses, which is a respectable-looking number.

"That is a 19 percent response rate" is the misreading worth killing. It is not a response rate. A response rate is answers divided by people selected, and nobody here was selected. The 1,300 silent students are not nonresponse; there was no sample for them to fail to respond to. The distinction decides the repair: nonresponse is attacked by chasing a known list of chosen individuals, and this poll has no such list, so there is nothing to chase and no way to patch the number after the fact.

It is also not a convenience sample, though both skip randomization. The difference is who did the choosing. In a convenience sample the researcher grabs whoever is easiest to reach; in a voluntary response sample the respondents put themselves forward.

Size is not the fix. An open poll with 100,000 answers is still self-assembled and still centered on whatever the motivated fraction believes, so the extra data only tightens the scatter around a wrong value. The repair is to select first: draw a random sample from a list of the population, then pursue the people you drew.

Full entry for voluntary response sample

Nonresponse bias

Collecting data and study design

Nonresponse bias occurs when people chosen for the sample do not respond, and those who do respond differ systematically from those who do not.

Nonresponse bias needs two conditions at once, and students usually check only the first. Selected people fail to give data, and those who responded differ from those who did not in a way that matters for the question. A low response rate by itself is not the bias, only the opening that lets it in: if the silent group would have answered like the responders, the estimate is not pushed anywhere.

A district emails a survey to a simple random sample of 250 teachers asking whether they have enough classroom supplies. 100 reply and 65 say yes, so the reported statistic is p^=65100=0.65\hat{p} = \frac{65}{100} = 0.65 (p-hat) on a response rate of 100250=0.40\frac{100}{250} = 0.40. Suppose that among the 150 who never replied, only 45 would have said yes, a rate of 0.30. The truth across the 250 selected teachers is 65+45250=0.44\frac{65 + 45}{250} = 0.44. The survey overstates supply adequacy by 0.21, and nothing in the returned data reveals it.

Here is the sentence that costs points: "only 100 responded, so treat it as a sample of 100." That describes the sample size correctly and the sample wrongly. The 100 are not a random subset of the 250; they are the subset that chose to reply, so a standard error built from n=100n = 100 describes scatter around 0.65, a center in the wrong place that a wider interval does not move. Emailing 1,000 teachers changes nothing: at the same reply rate the lean returns, measured more precisely.

Nonresponse happens after selection, and that is what separates it from its neighbors. All 250 teachers were on the list and could have answered. Had the list covered only full-time staff, part-timers would have had no chance of selection, which is undercoverage. Had all 250 replied while shading answers toward what an administrator wanted to hear, that would be response bias.

The defense is procedural, not statistical: follow-up contacts, a shorter instrument, and reporting the response rate next to the estimate. Topic 1.12 also wants a direction: the teachers who answered are plausibly those with the strongest feelings about supplies.

Full entry for nonresponse bias

Where each one fits in the course