Does a bigger sample fix bias? Why size fails
By Jude Wallis · Published
No. Bias is a lean built into the sampling method, so it repeats in the same direction on every sample and the size of the sample never touches it. A bigger sample shrinks sampling variability only. Fix bias by fixing the method: random selection, a complete list, a higher response rate.
AP Statistics: Unit 1 (topics 1.11 Random Sampling, 1.12 Potential Problems with Sampling). Bias versus sampling variability sits in Unit 1 topic 1.12 (Potential Problems with Sampling), with the random selection cure coming from topic 1.11, in a unit worth 20% to 30% of the multiple-choice section.
The short answer: bias lives in the method
Increasing the sample size does not reduce bias. Bias is a systematic error in how the sample is chosen, so it pushes the statistic in the same direction every single time the method is used. Collect twice as much data with a slanted method and you get twice as much slanted data.
Two separate things can go wrong with an estimate, and students collapse them into one.
- Bias: the method is centered on the wrong number. It is a property of the procedure, not of any one sample.
- Sampling variability: even a good method lands on a slightly different value each time, purely by chance.
Sample size is the dial for the second one only. Think of each sample as one dart. A bigger tightens the grouping of the darts. It does not move where you are aiming. If the sight is off by three inches, a perfectly tight group is three inches off, and a tighter group makes it look more convincing, not more correct.
That is the whole guide. Everything below is the arithmetic behind it and the wording the exam wants.
What a bigger sample actually buys you
Look at where appears in the sampling distribution. For a sample proportion from a random sample, the mean and standard deviation of are
and for a sample mean they are
Read the center formulas and the spread formulas separately. The center has no in it at all. The spread has , and only under a square root. Those spread formulas also need the sample to be under 10% of the population, or drawn with replacement, which is why the first worked example stops to check that before using them.
That square root sets the price of precision. To cut the standard deviation in half you need four times the data, because . To cut it to one tenth you need one hundred times the data. Sample size is expensive, and what you buy with it is a narrower spread around whatever the method is centered on.
Those center formulas, and , are exactly what people mean when they call a statistic unbiased: over many samples, it averages out to the parameter. But they hold because the sample is chosen at random. Randomness earns the correct center. Sample size never does. More on the machinery in sampling distributions explained.
Why n cannot move the center
Suppose a method can only reach part of the population. Call the proportion among the reachable people , and call the true population proportion . A random draw from the reachable group centers on , so the estimate is centered on no matter how many people you draw. The bias is
and there is no anywhere in that expression. Change the sample size from 50 to 5,000 and the two numbers on the right stay exactly where they were.
The classic demonstration is the 1936 Literary Digest presidential poll. Roughly 2.4 million people mailed back a ballot, one of the largest samples ever collected for a survey, and the poll predicted that Alf Landon would beat Franklin Roosevelt. Roosevelt won in a landslide. The magazine drew its names from lists that skewed toward people who could afford cars, telephones, and magazine subscriptions, and only a fraction of those contacted responded. Millions of responses did nothing about either problem. A modern poll of 1,000 people chosen at random beats it easily.
There is a genuine practical sense in which a bigger biased sample is worse, though the bias itself does not grow. A larger sample produces a narrower confidence interval, and a narrow interval around a wrong center is a confident wrong answer. Small biased samples at least look uncertain.
One honest footnote. Statisticians also use the word bias for a property of an estimator's formula, and a few of those do fade as grows: the sample standard deviation slightly underestimates on average, and that gap shrinks toward zero with more data. That is not what Unit 1 is asking about. Unit 1 bias is about who ends up in the sample, and no sample size repairs it.
The four named biases, and what each one needs
The course names four kinds of bias, and how to identify the type of bias walks through pinning the right name on a scenario. None of them is a sample size problem, and each one has a different repair.
| Type of bias | Where the lean comes from | Does a bigger sample help? | What actually helps |
|---|---|---|---|
| Voluntary response | People opt themselves in, and volunteers tend to hold stronger opinions | No | Select the sample yourself, at random |
| Undercoverage | Part of the population cannot be selected, or is less likely to be | No | Repair the list you sample from |
| Nonresponse | Selected people do not answer, and non-responders differ from responders | No | Raise the response rate, follow up with the people you missed |
| Response bias | Answers lean away from the truth, from leading wording or a sensitive question | No | Reword the question, guarantee anonymity, train interviewers |
Run the middle column down the table. A larger online poll still collects volunteers. A larger sample drawn from a landline directory still misses everyone without a landline. A larger mailing with a 20% response rate still hears from the same self-selected fifth. A larger sample answering the same leading question still gets the same slanted answers. In every case the flaw is in the recipe, and doubling the batch keeps the flaw.
What actually reduces bias
Fix the method, in this order.
- Use random selection. Chance decides who is in, not the researcher and not the respondent. This is the single biggest defense, because it removes the human choices that create the lean. The standard designs are compared in simple random vs stratified sampling.
- Fix the sampling frame. The list you sample from should cover the whole population you want to describe. Undercoverage happens before any randomness is applied, so no amount of good random sampling from a bad list saves you.
- Chase the nonresponders. Follow-up calls, second mailings, and incentives raise the response rate. A 70% response rate leaves far less room for nonresponse bias than a 20% response rate on ten times as many people.
- Write neutral questions. Ask a plain question, keep the wording out of the way, and let sensitive answers be anonymous.
Once the method is sound, sample size finally does something useful. An unbiased method centered on the right number gets more precise as grows, which is why the sample size calculator is worth using at the planning stage. The order matters: get unbiased first, get precise second. Precision applied to a biased method just sharpens the error.
Also worth knowing: a census does not automatically escape this. Measuring everyone removes sampling variability, since there is no sampling left to vary, but it can still suffer nonresponse and response bias. See census vs sample survey.
How the AP exam tests this
This sits in topic 1.12, Potential Problems with Sampling, inside Unit 1, which is worth 20% to 30% of the multiple-choice section. The full topic page is 1.12 Potential Problems with Sampling.
On multiple choice, the trap answer is usually the one that reads "increase the sample size" as the fix for a described bias. The correct choice names a change to the method. A second trap runs the other way: a question describes a method with no flaw beyond a small sample, and the fix really is more data, because the problem there is variability, not bias.
On free response, "take a larger sample" earns nothing when the scenario describes bias, and graders see it constantly. Write two things instead.
- Name the bias and say where in the procedure it enters, in the context of the study.
- Say which direction it pushes the estimate, and why.
A usable sentence, for a scenario where students were invited to comment on a new parking fee: "Because only students angry about the fee bothered to respond, the sample over-represents opponents, so the reported proportion of students who oppose the fee is too high. A larger sample of volunteers would repeat the same lean, so the fix is to select respondents at random instead."
The vocabulary that keeps these answers straight is in parameter vs statistic, and the general scoring habits are in the AP Statistics FRQ guide.
Four times the data, the same wrong center
A university has 10,000 students: 6,000 underclassmen, of whom 10% own a car, and 4,000 upperclassmen, of whom 60% own a car. A reporter wants the proportion of all students who own a car and samples at random from an online student directory that, unknown to the reporter, lists only the 4,000 upperclassmen, so no underclassman can be selected. (a) Find the true proportion of all students who own a car. (b) Find the center and standard deviation of the reporter's sample proportion for and for . (c) Say what changed.
Count the car owners in the whole population. Underclassmen: . Upperclassmen: . Total owners .
True parameter: , so 30% of all students own a car.
The method can only reach upperclassmen, and 60% of them own a car. A random sample from that group is centered on for every sample size.
Bias , an overestimate of 30 percentage points.
Standard deviation at : .
Standard deviation at : .
Check the 10% condition before trusting those formulas: 50 and 200 are 1.25% and 5% of the 4,000 upperclassmen, both well under 10%.
Compare. The sample size went up by a factor of 4 and the standard deviation fell by a factor of , from 0.0693 to 0.0346. The center did not move: still 0.60.
Push it one step. Suppose the sample gives . A 95% interval has margin of error , giving 0.532 to 0.668. The true 0.30 is not close to that interval, and a larger sample would only narrow it further.
. The reporter's method centers on 0.60 at every sample size, a bias of . The standard deviation drops from 0.069 at to 0.035 at . Four times the data halved the spread and left the bias untouched, so the bigger sample is a more precise wrong answer.
Ten times the mailing, the same nonresponse lean
A city of 200,000 residents is surveyed by mail about a new tax. 40% of residents follow local politics closely; the other 60% do not. Among the close followers, 60% return the survey and 80% support the tax. Among the rest, 20% return the survey and 30% support the tax. Addresses are drawn at random, so a mailing of any size is expected to split 40/60 the same way. (a) Find the true level of support. (b) Using expected counts, find the value the estimate from the returns is centered on when 1,000 surveys are mailed. (c) Repeat with 10,000 mailed, and say what did and did not change.
True support is the weighted average across the two groups: , so 50% of residents support the tax.
Mail 1,000 surveys. The mailing is expected to reach close followers and others.
Expected returns: close followers ; others . About 360 come back, a response rate of 36%.
Expected supporters among those returns: and , so 228.
The estimate from the returns is centered on , about 63.3%. Bias , an overestimate of roughly 13 percentage points.
Now mail 10,000 surveys. Expected reach: 4,000 close followers and 6,000 others. Expected returns: and , so about 3,600 come back.
Expected supporters: and , so 2,280.
The estimate is now centered on , about 63.3%, the same center as the small mailing.
See why. Every expected count multiplied by 10, so the ratio could not change. The response rates and the support rates are what set the center, and mailing more envelopes changes neither.
The two mailings are not identical, though. Any one mailing lands near its center rather than on it, and the bigger mailing lands nearer: against . Ten times the mail buys a spread about three times tighter, around the same wrong 63.3%.
True support is 50%. The 1,000-piece mailing is expected to give 228 supporters out of 360 returns and the 10,000-piece mailing 2,280 out of 3,600, both centered on 63.3%. Ten times the mail leaves the same 13.3 point overestimate and only tightens how closely a single mailing clusters around it. Moving the center takes a higher response rate or follow-up with the people who did not reply.
Frequently asked questions
Does increasing sample size reduce bias?
No. Bias comes from how individuals are selected, so it repeats in the same direction in every sample the method produces. A larger sample reduces sampling variability, which is the chance-to-chance wobble, and leaves the bias exactly where it was. Only a change to the method reduces bias.
Can a small sample be unbiased?
Yes. A simple random sample of 20 people is unbiased: over many such samples the statistic averages out to the parameter. It is just very imprecise, since the standard deviation of the statistic is large when n is small. Unbiased and precise are two separate properties, and sample size only controls the second.
Does a bigger biased sample make things worse?
The bias itself does not grow, but the reported uncertainty shrinks around the wrong center. A large biased sample produces a narrow confidence interval that can miss the parameter completely, so it presents a wrong answer with high confidence. A small biased sample is wrong too, and at least looks uncertain.
Does surveying everyone, a census, remove bias?
Not automatically. A census removes sampling variability, since nothing is left to sample. It can still carry response bias from bad question wording and nonresponse bias from the people who never reply, and those are the same problems that a bigger sample fails to fix.