Question Wording Bias vs Response Bias

Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.

Question wording bias

Collecting data and study design

Question wording bias occurs when a confusing or leading survey question pushes responses away from the true value in one direction.

Question wording bias sits in the instrument, not in the sample. The right people were selected and every one of them answered; the answers lean because of how the question was put. That makes it one kind of response bias, and its signature is that two wordings of the same underlying question, put to two equivalent random samples, come back with different answers.

Here are two versions of one question about the same school policy.

  • Do you support a ban on phone use during class, so that students are not distracted during lessons?
  • Do you support a ban on phone use during class, even though it would stop students reaching a parent in an emergency?

Split 800 students at random, 400 to each version. Suppose the first returns p^=0.61\hat{p} = 0.61 in favor and the second 0.380.38. Each survey's own 95 percent margin of error is about 0.048, so the gap of 0.23 is roughly five times either one, and the standard error of the difference is 0.034, putting the gap almost 7 standard errors out. Chance did not produce that. The neutral version, "do you support a ban on phone use during class?", is the one that estimates anything.

The wrong diagnosis is "the sample was not representative," or "they should have surveyed more students." Everybody selected answered, so nobody is missing, and asking the loaded question of 4,000 students returns the same lean with a narrower interval around it. The lean lives in the instrument, so only the instrument fixes it.

Not every wording problem is this one. A double-barreled question, such as asking whether people support raising taxes to fund parks and libraries, produces answers nobody can interpret, since a yes and a no each cover two positions; that is confusion rather than a push in one direction. Question order pushes in a direction, and a sensitive question asked face to face pulls toward the socially acceptable answer whatever words it uses.

Potential problems with sampling is topic 1.12 in Unit 1.

Full entry for question wording bias

Response bias

Collecting data and study design

Response bias occurs when respondents give systematically inaccurate answers, for example due to confusing question wording or pressure to answer a certain way.

Response bias lives in the answers, not in the sample. The right people were reached, they responded, and what they said leans away from the truth in a consistent direction. The usual causes are leading or confusing wording, a sensitive question asked without anonymity, an interviewer whose presence changes what people will say, and faulty recall. Because the lean sits in the instrument, a flawless random sample with a 100 percent response rate can still return a badly wrong number.

A coach asks each of her 200 athletes face to face whether they skipped a scheduled workout last month, and 18 say yes, a reported rate of 18200=0.09\frac{18}{200} = 0.09. The same 200 athletes then answer the same question on an anonymous form, and 62 say yes, or 62200=0.31\frac{62}{200} = 0.31. The gap of 0.22 cannot be sampling variability, because it is the same 200 people both times. It is a property of how the question was asked.

"The athletes lied, so the data are just noisy" packs two errors into one clause. Noise is symmetric and averages out across samples, while under-reporting an embarrassing behavior pushes every repetition of that survey the same way, which is exactly what makes it bias. Blaming the respondents also misplaces the repair. What differed between 0.09 and 0.31 was the format of the question, so that is where the fix belongs: neutral wording, anonymity, or a measurement that does not rely on self-report at all.

The direction is usually predictable. Socially approved behavior gets over-reported and disapproved behavior under-reported, so self-reported exercise and voting run high while self-reported drinking and cheating run low. A leading question pushes answers toward whatever it hints at. Naming which way, in the context of the study, is the half of the answer graders look for.

Two things response bias is not. It is not nonresponse bias, because everyone here answered. It is not undercoverage, because everyone here was reachable. It also survives a perfect sampling design, since fixing who ends up in the sample does nothing to the question those people are handed.

Full entry for response bias

Where each one fits in the course