Undercoverage vs nonresponse: list or no reply

By Jude Wallis · Published

Undercoverage means the person was never on the list the sample was drawn from, so their chance of selection was zero. Nonresponse means they were selected and gave no answer. Either can bias the estimate, but only nonresponse leaves a number in your own records: a response rate.

AP Statistics: Unit 1 (topics 1.12 Potential Problems with Sampling). Undercoverage and nonresponse are both sampling problems, and topic 1.12 of the Fall 2026 AP Statistics course is titled Potential Problems with Sampling. It sits in Unit 1, which is 20% to 30% of the multiple-choice section.

Undercoverage vs nonresponse: the short answer

Both leave people missing from the data, and the difference is when they went missing. Undercoverage is a failure of the sampling frame, the list your method actually draws from: someone in the population is not on it, so their chance of being selected is zero. Nonresponse bias is a failure after selection: the person was on the list, was picked, and never gave you an answer.

Draw the timeline once and the pair stops blurring. Undercoverage happens before anyone is selected. Nonresponse happens after selection, whether or not the person was ever reached. Three things they do share: both are systematic, both survive a larger sample, and both push the estimate only when the missing people differ on the variable being measured.

The consequence worth carrying out of this page is the one that rarely gets said. A study can measure its own nonresponse and cannot measure its own undercoverage. The response rate is a number the study computes from its own records. Coverage has no matching number, because everything the study holds came from the frame.

Undercoverage: the failure is in the list

Three groups sit inside one another. The population is who the question is about. The frame is the list, roster, or directory the method can actually draw from. The sample is who gets selected out of the frame. Undercoverage is the gap between the first two.

A pollster who dials numbers from a landline directory has a frame of landline households. Cell-only households are in the population and not in the frame, so their selection probability is exactly zero. Not small. Zero, in this sample and in every sample the method will ever produce. That is why the instinct to survey more people is the wrong repair: enlarging the sample draws harder on the same frame, so the estimate stays centered on the frame's value while its standard error shrinks. More data makes a wrong number look authoritative.

An incomplete frame is not automatically a biased estimate, though. If the excluded households would have answered like the covered ones, the estimate sits on the parameter anyway. The same directory that is harmless for a question about household size can wreck a question about a transit measure. So the description that earns credit has two halves: name who was left off the list, then say how those people differ on the variable being measured.

Nonresponse bias: the failure is after selection

Here the frame did its job. The person was on the list, the random draw picked them, and no data came back. They refused, nobody was home, the email went unopened, or they skipped that one item on the form.

Nonresponse bias needs two conditions at once, and only the first is visible. Some selected people give no data, and those people differ from the ones who answered in a way that matters for the question. A 40% response rate in which the silent group would have answered exactly like the repliers costs precision, not accuracy. It is the opening, not the bias.

What makes the second condition likely is that the repliers chose themselves. Random selection decided who got the form; each person decided whether to send it back. So the returned forms are not a random subset of the people selected, they are the subset willing to answer, and a bigger mailing at the same reply rate simply collects more of the same willing people. That is also why the repair is contact rather than volume: the people you are missing are named, on your list, and reachable.

The differences side by side

FeatureUndercoverageNonresponse bias
When it happensBefore selection, when the frame is builtAfter selection, when answers come back
Chance the person was selectedZeroNormal; they were selected
Whose failure it isThe listThe reply
Visible inside the studyNoYes, as the response rate
A bigger sampleDoes not helpDoes not help
Follow-up contactsDo not helpShrink it, by reaching the silent
Counts as selection biasYesNo, selection worked as designed

The first row and the fourth are the ones to memorize. Every other difference on this page follows from where in the process the failure sits.

Only one of them shows up in your own data

Every survey that selects a sample can compute one quality number without leaving its own records: the response rate, replies divided by people selected. Both counts are things the study did itself, so nonresponse announces its own size. A study that mails 400 forms and gets 160 back knows it lost 240 people, knows which 240, and can go after them.

Coverage has no such number. Every record a study holds came from the frame, so the data cannot report on people the frame never listed. Nothing is missing from the spreadsheet, because those rows were never scheduled to exist. A survey that covers 70% of the population and hears back from every single person it contacts looks flawless from inside: a 100% response rate, no refusals, clean data, and an estimate aimed at the wrong value. The worked example below puts numbers on exactly that.

To put a number on undercoverage you have to leave the study. Compare the frame against something outside it: an enrollment count, a payroll list, a published population total. That comparison is the only way to size the gap. Noticing that a gap is there at all needs nothing from outside, because the problem itself describes how the frame was built, which is why the first question to put to a survey is not how many answered but who the method could never have reached.

Two things follow. A response rate is a poor summary of survey quality, because it measures one failure precisely and is blind to the other completely. And the reason a larger nn fixes neither is the same in both cases: sample size moves the spread of the sampling distribution of p^\hat{p}, never its center, the point worked through in does a bigger sample fix bias.

The classic mix-up and how to avoid it

One question separates them: could that person have been selected at all? No means undercoverage. Yes, and they gave nothing, means nonresponse. How to tell which type of bias a survey has runs that question inside a four-step flow covering the whole family of sampling biases, which is the version to use under exam pressure.

What trips people is which sentence of the problem to read. Nonresponse is loud. It arrives as arithmetic, "only 34 of the 100 replied", and it pulls the eye straight to it. Undercoverage is quiet. It hides in a clause about how the list was built: "numbers were drawn from the landline directory", "the survey was emailed to everyone in the staff directory", "the roster was printed on Monday". Read the frame sentence first, because a scenario can carry both problems and only one of them announces itself.

Two more habits keep the pair apart. Undercoverage is a species of selection bias, alongside convenience and voluntary response samples, and what is selection bias is where that family lives; nonresponse is not selection bias, because selection worked exactly as designed. And neither is established by a missing group alone. A landline frame that misses a third of a town estimates the town's mean household size fine and its opinion of a transit measure badly, and the same is true of the third of a sample that never wrote back. Practice naming both in scenarios at bias practice.

A 100% response rate that misses by more than a 40% one

A company has 2,000 employees: 1,400 at the main site and 600 at a satellite warehouse whose staff have no internal email accounts. Management wants the proportion who would use a new commuter shuttle. In truth 30% of main-site employees and 60% of warehouse employees would use it. Survey A emails everyone in the internal email directory, and all of them reply. Survey B mails a form to a simple random sample of 400 drawn from the full payroll list, and 160 forms come back. Which survey is closer to the truth, and which one warns you that it is wrong?

  1. Compute the parameter. Would-use employees =1400(0.30)+600(0.60)=420+360=780= 1400(0.30) + 600(0.60) = 420 + 360 = 780, so p=7802000=0.39p = \frac{780}{2000} = 0.39.

  2. Survey A: identify the frame. The internal email directory holds the 1,400 main-site employees and none of the 600 warehouse staff, so 30% of the population has selection probability zero. That is undercoverage.

  3. Survey A: compute what it reports. All 1,400 reply, and 1400(0.30)=4201400(0.30) = 420 of them would use the shuttle, so p^=4201400=0.30\hat{p} = \frac{420}{1400} = 0.30 (p-hat).

  4. Survey A: compute its response rate, 14001400=1.00\frac{1400}{1400} = 1.00. The only quality number the study can compute from its own records is a perfect 100%, and the estimate is still 0.390.30=0.090.39 - 0.30 = 0.09 too low, 9 percentage points.

  5. Survey B: check coverage. The payroll list holds all 2,000 employees, so every one of them could be selected. No undercoverage. In expectation the 400 selected split proportionally: 400×14002000=280400 \times \frac{1400}{2000} = 280 from the main site and 400×6002000=120400 \times \frac{600}{2000} = 120 from the warehouse.

  6. Survey B: account for the 160 replies. Suppose main-site employees return the form at 25% and warehouse employees, who care more about the shuttle, at 75%. Then 280(0.25)=70280(0.25) = 70 and 120(0.75)=90120(0.75) = 90 reply, for 70+90=16070 + 90 = 160 forms and a response rate of 160400=0.40\frac{160}{400} = 0.40.

  7. Survey B: compute what it reports. Among repliers, 70(0.30)=2170(0.30) = 21 main-site and 90(0.60)=5490(0.60) = 54 warehouse employees would use the shuttle, so p^=21+54160=75160=0.46875\hat{p} = \frac{21 + 54}{160} = \frac{75}{160} = 0.46875. Check: 0.46875×160=750.46875 \times 160 = 75.

  8. Survey B: measure the error. 0.468750.39=0.078750.46875 - 0.39 = 0.07875, about 7.9 percentage points too high. That is nonresponse bias, since the warehouse staff who replied are over-represented relative to their share of the 400 selected.

  9. Compare. Survey A is off by 9 points with a spotless response rate. Survey B is off by about 7.9 points with a response rate that shouts. The survey that looks worse from inside is the more accurate of the two.

  10. Compare the repairs. Survey B knows the identity of all 400160=240400 - 160 = 240 silent employees and can mail them again. Survey A holds nothing about the 600 warehouse staff; only the payroll count, which lives outside the survey, reveals that they exist.

Survey A reports 0.30 against a true 0.39, an undercoverage bias of 9 percentage points hidden behind a 100% response rate. Survey B reports 0.46875, about 7.9 percentage points high, and its 40% response rate flags the problem. Nonresponse is the smaller error here and the only one visible from inside the study.

Frequently asked questions

Does a high response rate mean a survey is unbiased?

No. The response rate only describes the people who were selected, so it says nothing about who the frame left out. A survey drawn from a list covering 70% of the population can hit a 100% response rate and still be aimed at the wrong value, which is the worked example above.

How do you detect undercoverage?

Not from the survey data, because every record came from the frame. You compare the frame against a description of the population from outside the study: an enrollment count, a payroll list, a published total. If the frame is missing a group, name it and say how that group differs on the variable measured.

Is undercoverage a type of selection bias, and is nonresponse?

Undercoverage is selection bias: the method gave some people no chance of being chosen. Nonresponse is not, because selection worked as designed and the lean entered one stage later, when some selected people answered and others did not.

Which one can follow-up contacts fix?

Nonresponse only. The silent people are on your list, named and reachable, so a second mailing or a phone call brings some of them in. Undercoverage cannot be chased that way, since the people you need were never on the list. Only a frame that covers the population reaches them.

Can there be missing people without any bias?

Yes, for both. Bias needs the missing group to differ on the variable being measured. A landline frame that skips cell-only households estimates mean household size fine and support for a transit measure badly, and a 40% response rate is harmless if the silent people would have answered like the repliers.