Confounding Variable vs Response Bias
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Confounding variable
Collecting data and study design
A confounding variable is associated with both the explanatory variable and the response, so its effect and the explanatory variable's cannot be told apart.
Confounding is a two-part test and a variable has to pass both parts. The candidate must be associated with the explanatory variable, so the groups being compared differ on it, and it must also be associated with the response. Fail either part and the variable is not a confounder.
A district finds that students who eat the free school breakfast score higher on a reading test than students who do not. Explanatory variable: eats the school breakfast. Response: reading score. Family income passes both parts, since lower-income families take up a free breakfast at higher rates and income is separately related to reading scores, so the breakfast gap and the income gap are the same gap seen twice. Nothing in these data says which one moved the score.
Now a variable that fails. Suppose 12 percent of the breakfast eaters are left-handed and 12 percent of the non-eaters are as well. Handedness is spread evenly across the two groups, so it is not associated with the explanatory variable, and whatever it does to reading it cannot account for a difference between them. That is the correction worth keeping, because the wrong version is everywhere: "any variable that could affect the reading score is a confounding variable." No. A variable that affects the response but has no tie to who ate breakfast is an extraneous variable. It adds spread to the scores rather than a lean to the comparison.
The other half of the test fails just as often. Riding the school bus is strongly associated with eating the school breakfast, because bus riders arrive early, but if riders and walkers read alike then bus riding explains none of the gap.
Confounding is a feature of how a study was built, not something you can spot in the numbers. Random assignment attacks the first link by making the treatment groups similar on average on every other variable, measured or not, though in a small experiment chance can still leave a group tilted. Topic 1.13 Experimental Design is where this sits.
Response bias
Collecting data and study design
Response bias occurs when respondents give systematically inaccurate answers, for example due to confusing question wording or pressure to answer a certain way.
Response bias lives in the answers, not in the sample. The right people were reached, they responded, and what they said leans away from the truth in a consistent direction. The usual causes are leading or confusing wording, a sensitive question asked without anonymity, an interviewer whose presence changes what people will say, and faulty recall. Because the lean sits in the instrument, a flawless random sample with a 100 percent response rate can still return a badly wrong number.
A coach asks each of her 200 athletes face to face whether they skipped a scheduled workout last month, and 18 say yes, a reported rate of . The same 200 athletes then answer the same question on an anonymous form, and 62 say yes, or . The gap of 0.22 cannot be sampling variability, because it is the same 200 people both times. It is a property of how the question was asked.
"The athletes lied, so the data are just noisy" packs two errors into one clause. Noise is symmetric and averages out across samples, while under-reporting an embarrassing behavior pushes every repetition of that survey the same way, which is exactly what makes it bias. Blaming the respondents also misplaces the repair. What differed between 0.09 and 0.31 was the format of the question, so that is where the fix belongs: neutral wording, anonymity, or a measurement that does not rely on self-report at all.
The direction is usually predictable. Socially approved behavior gets over-reported and disapproved behavior under-reported, so self-reported exercise and voting run high while self-reported drinking and cheating run low. A leading question pushes answers toward whatever it hints at. Naming which way, in the context of the study, is the half of the answer graders look for.
Two things response bias is not. It is not nonresponse bias, because everyone here answered. It is not undercoverage, because everyone here was reachable. It also survives a perfect sampling design, since fixing who ends up in the sample does nothing to the question those people are handed.