Identifying bias practice problems (8 solved)

By Jude Wallis · Updated

This set has 8 problems on naming the bias in a described study, predicting which way the estimate is pushed, and separating bias from sampling variability. Every answer names the mechanism and the direction, because naming the type alone is half of what the question asks for.

AP Statistics: Unit 1 (topics 1.11 Random Sampling, 1.12 Potential Problems with Sampling). These problems cover topic 1.12 (Potential Problems with Sampling) and the random selection repair in topic 1.11, both in Unit 1 of the Fall 2026 AP Statistics course, where Unit 1 carries the heaviest multiple-choice weight at 20 to 30%.

What these problems build

These 8 problems build one habit: find the stage where the lean entered, name it with the course's vocabulary, then say which way the estimate is pushed and why. Bias in a sampling method is a systematic error that leaves the statistic consistently larger or consistently smaller than the parameter it estimates. That is a property of the procedure, not of any one sample, which is why you diagnose it from the description of the method rather than from the number it produced.

Four named problems, each entering at a different stage.

  • [Undercoverage](/glossary/undercoverage) enters at the list: part of the population was left off it, or the method makes that part less likely to be picked.
  • [Voluntary response](/glossary/voluntary-response-sample) enters at selection: the people in the sample put themselves there by answering an open invitation.
  • [Nonresponse](/glossary/nonresponse-bias) enters after selection: the right people were chosen and some of them never answered.
  • [Response bias](/glossary/response-bias) enters at the data: the right people answered, but the answers lean away from the truth. Leading or confusing question wording and self-reported behavior are the usual causes.

Naming the type is half of an answer. The other half is a direction attached to the actual variable, and every solution here writes it in the same shape: because the group that is over- or under-represented tends to differ in some specific way, the method overstates or understates the parameter, in context. The arithmetic is never there for its own sake. It exists to pin down a direction, bound how far off an estimate could be, or show that a sample proportion p^\hat{p} (p-hat) tightens around the wrong number as the sample size nn grows.

For the full walkthroughs, read how to identify the type of bias and does a bigger sample fix bias, which problems 3, 4, 7, and 8 all circle back to. Once you can name the flaw, can you generalize results covers what a damaged study still supports. The course topic is 1.12 Potential Problems with Sampling, and the designs that avoid these problems in the first place are worked in sampling methods practice.

Problem 1

For each study, name the type of bias and say whether the estimate is pushed too high or too low.

(a) A bakery wants the proportion of all its customers who would buy a new sourdough loaf. It leaves a stack of comment cards by the register inviting customers to write in if they would buy the new loaf, and counts the cards dropped in the box.

(b) A library wants the mean number of books its 8,000 cardholders read last year. It draws a random sample of 300 names from the list of cardholders who have visited in the past 30 days.

(c) A university mails a survey to a simple random sample of 800 alumni asking whether they gave money to any charity last year. 210 surveys come back.

(d) A school nurse asks each of 60 randomly selected students, face to face in her office, how many energy drinks they had last week.

(e) A town puts this item on a mailed ballot: "Should the town waste 4 million dollars on a new roundabout?"

Show the worked solution
  1. (a) Who put these people in the sample? The customers did, by choosing to fill out a card. That is voluntary response bias. Customers with a strong feeling about the bakery's bread are the ones who bother, and interest in a new loaf is exactly what is being measured, so the estimate of the proportion who would buy overstates the truth.

  2. (b) Could every cardholder have been selected? No. The list is recent visitors only, so any cardholder who has not been in for a month had no chance of selection at all. That is undercoverage. The people left off the list are the light users, so the mean number of books read comes out too high.

  3. (c) The 800 were selected properly by a random mechanism, and 590 of them never answered, so this is nonresponse bias, not undercoverage. Two things push the same way here: an alum who gives to charity has a creditable answer to write down, and an alum who gives to this university in particular is likelier to open its mail in the first place. Both raise the return rate among givers, so the estimated proportion who gave is too high.

  4. (d) These 60 were randomly selected, and they all answered, so the selection is sound and the lean is in the answers themselves. That is response bias, from a self-reported behavior that adults disapprove of, asked face to face by a school employee. Students shade their counts downward, so the mean number of energy drinks is too low.

  5. (e) Everyone on the mailing list could be selected and the answers come from the ballot itself, so the problem is again response bias, in its question wording form. The word "waste" supplies a verdict before the question is asked, and it sits on one side only. Support for the roundabout is understated.

  6. Sort the five. Selection created the problem in (a) and (b), the gap between selection and data created it in (c), and the data itself carries it in (d) and (e). Three estimates are pushed up and two are pushed down.

(a) Voluntary response, too high: only customers who care about the bakery's bread fill out a card. (b) Undercoverage of inactive cardholders, too high: the missing people are the light readers. (c) Nonresponse, too high: donors are likelier to return a survey about giving. (d) Response bias from face-to-face self-report, too low: students underreport a disapproved behavior. (e) Response bias from question wording, too low: "waste" pushes voters away from the roundabout.

Problem 2

A gym chain wants the mean number of days per week its 3,000 members visit. A staffer stands at the door one Tuesday evening and surveys every 5th person who walks in, until 60 members have answered.

Suppose the truth is that 900 of the 3,000 members come 4 days a week and the other 2,100 come 1 day a week, and that each member's visits are spread evenly across the week.

(a) Name the bias in this method and explain the mechanism. (b) Find the true mean number of visits per week. (c) Of the visits made in a week, what fraction are made by the frequent members? Use that to find the mean weekly visits of the member the staffer actually intercepts. (d) State the size and direction of the bias.

Show the worked solution
  1. (a) Name it. Taking every 5th arrival is a systematic selection, and that part is fine. The damage is in who can be at the door at all. A member who never comes on a Tuesday evening cannot be selected, and a member who comes 4 days a week is 4 times as likely to be there as a member who comes once. That is undercoverage: the method makes part of the population less likely to be selected, and it does so based on the very variable being measured. It is also a convenience sample, since the staffer surveys whoever happens to walk past.

  2. (b) True mean. Add up all the weekly visits and divide by the number of members: xˉ=900(4)+2100(1)3000=3600+21003000=57003000=1.9\bar{x} = \frac{900(4) + 2100(1)}{3000} = \frac{3600 + 2100}{3000} = \frac{5700}{3000} = 1.9 days per week, where xˉ\bar{x} (x-bar) is the mean visits per member.

  3. (c) Who walks through the door. In one week the frequent members make 900×4=3600900 \times 4 = 3600 visits and the occasional members make 2100×1=21002100 \times 1 = 2100 visits, so 3600+2100=57003600 + 2100 = 5700 visits happen in total. Because visits are spread evenly across the week, treat the person the staffer intercepts as one randomly chosen visit. Then 360057000.6316\frac{3600}{5700} \approx 0.6316 of the people at the door are frequent members, even though frequent members are only 9003000=0.30\frac{900}{3000} = 0.30 of the membership.

  4. Average over the visits, not over the members. The mean weekly visits of an intercepted member is 3600(4)+2100(1)5700=14400+21005700=1650057002.8947\frac{3600(4) + 2100(1)}{5700} = \frac{14400 + 2100}{5700} = \frac{16500}{5700} \approx 2.8947, or about 2.89 days per week.

  5. (d) Size and direction. The method is centered on 2.89 while the parameter is 1.9, so the bias is 2.89471.90.992.8947 - 1.9 \approx 0.99, about one full day per week too high. The direction follows directly from the mechanism: the people who use the gym most are the people most likely to be standing in the doorway, so the sample overrepresents them and the estimate overstates how often the typical member comes.

(a) Undercoverage, in a convenience sample: a member who visits 4 days a week is 4 times as likely to be at the door as a member who visits once, so selection probability rises with the variable being measured. (b) 57003000=1.9\frac{5700}{3000} = 1.9 days. (c) 360057000.6316\frac{3600}{5700} \approx 0.6316 of visits are made by frequent members, giving an intercepted mean of 1650057002.89\frac{16500}{5700} \approx 2.89 days. (d) Bias 2.891.9=0.99\approx 2.89 - 1.9 = 0.99 days too high, because heavy users are overrepresented at the door.

Problem 3

An online retailer emails a one-question satisfaction survey to all 40,000 of its customers. 5,000 answer, and 4,300 of them say they are satisfied. Customers who were disappointed had mostly stopped opening the retailer's email months earlier.

(a) Find p^\hat{p} and the response rate. (b) Name the bias, the mechanism, and the direction. (c) A later audit that tracks down a random subsample of the non-answerers puts true satisfaction across all 40,000 customers at 0.71. How large is the bias? (d) A manager says the fix is to email 120,000 customers next quarter. If the response rate holds at the same level, find the standard deviation of p^\hat{p} before and after, and compare each one to the bias. (e) Answer the manager.

Show the worked solution
  1. (a) Two rates, two meanings. The sample proportion is p^=43005000=0.86\hat{p} = \frac{4300}{5000} = 0.86, and the response rate is 500040000=0.125\frac{5000}{40000} = 0.125, or 12.5%. The second number is the warning sign: 35,000 selected customers are missing from the result.

  2. (b) Name it. The right people were contacted and most never replied, so this is nonresponse bias. The mechanism is stated in the setup: staying subscribed to the retailer's email is itself a mild sign of satisfaction, so the 5,000 who opened and answered lean happier than the 35,000 who did not. The estimate overstates satisfaction.

  3. (c) Size of the bias. The method is centered on 0.86 and the parameter is 0.71, so the bias is 0.860.71=0.150.86 - 0.71 = 0.15, or 15 percentage points too high.

  4. (d) What the extra email buys. Holding the response rate at 12.5% gives 120000×0.125=15000120000 \times 0.125 = 15000 answers. Treating the answerers as a random draw from the answering group, whose satisfaction rate is about 0.86, the standard deviation of p^\hat{p} is 0.86(0.14)5000=0.12045000=0.000024080.0049\sqrt{\frac{0.86(0.14)}{5000}} = \sqrt{\frac{0.1204}{5000}} = \sqrt{0.00002408} \approx 0.0049 before, and 0.120415000=0.00000802670.0028\sqrt{\frac{0.1204}{15000}} = \sqrt{0.0000080267} \approx 0.0028 after. Tripling nn divides the spread by 31.732\sqrt{3} \approx 1.732, so it falls by about 0.2 percentage points.

  5. Compare the two quantities. Before, the bias is 0.150.004931\frac{0.15}{0.0049} \approx 31 standard deviations. After, it is 0.150.002853\frac{0.15}{0.0028} \approx 53 standard deviations. The extra email makes the estimate more precisely wrong: the center never moved, and the noise that used to partly disguise the error got smaller.

  6. (e) The answer to give. Sample size is the dial for sampling variability only, and this estimate's problem is a 15 point lean that lives in the method. Raise the response rate instead: follow up by phone or text with a random subsample of the customers who did not answer, and report the estimate from that subsample, which represents the missing 87.5% rather than ignoring them.

(a) p^=43005000=0.86\hat{p} = \frac{4300}{5000} = 0.86, response rate 500040000=0.125\frac{5000}{40000} = 0.125. (b) Nonresponse bias: unhappy customers had already stopped opening the retailer's email, so the answerers lean satisfied and 0.86 is too high. (c) Bias =0.860.71=0.15= 0.86 - 0.71 = 0.15, 15 percentage points too high. (d) The standard deviation falls from about 0.0049 to about 0.0028, roughly 0.2 percentage points, while the bias stays at 15 points, moving from about 31 to about 53 standard deviations. (e) Emailing more people cannot move the center; chase the non-answerers with a random follow-up subsample instead.

Problem 4

A teacher knows that exactly 0.24 of her school's 5,000 students bike to school. Her class simulates two survey methods, running each one 200 times and recording the 200 sample proportions.

Method A: a simple random sample of 50 students from the full enrollment roster. The 200 results have mean 0.241 and standard deviation 0.060.

Method B: a random sample of 50 students drawn from the roster of the 2,400 students who live within one mile of the school. The 200 results have mean 0.385 and standard deviation 0.069.

(a) Which method is biased, and which two numbers show it? (b) Estimate the size of the bias and explain the mechanism and direction. (c) Method A missed 0.24 by 0.001. Is that evidence of bias? Support the answer with a calculation. (d) Both classes now repeat the simulation with samples of 200 instead of 50. Predict all four summary numbers. (e) State the difference between bias and sampling variability in one sentence.

Show the worked solution
  1. (a) Read the centers, not the spreads. Bias is a question about where a method is centered, so compare each mean to the parameter p=0.24p = 0.24. Method A centers at 0.241 and Method B centers at 0.385. Method B is the biased one, and the two numbers that show it are its mean of 0.385 against the true 0.24.

  2. (b) Size, mechanism, direction. The bias is about 0.3850.240=0.1450.385 - 0.240 = 0.145, or 14.5 percentage points. The mechanism is undercoverage: students living more than a mile away were never on Method B's list. Distance is tied to the variable being measured, since a student two miles out is far likelier to ride the bus than to bike, so the reachable group bikes more than the school does. The estimate overstates the proportion who bike.

  3. (c) Is 0.001 evidence of anything? No, and here is the check. The mean of 200 independent sample proportions has standard deviation 0.0602000.06014.1420.0042\frac{0.060}{\sqrt{200}} \approx \frac{0.060}{14.142} \approx 0.0042, so a miss of 0.001 is about a quarter of one standard deviation. That is ordinary chance. Method B's miss, by contrast, is 0.1450.069/2000.1450.004930\frac{0.145}{0.069 / \sqrt{200}} \approx \frac{0.145}{0.0049} \approx 30 standard deviations, which chance does not produce.

  4. (d) Predict the four numbers. Quadrupling the sample size from 50 to 200 divides the standard deviation of p^\hat{p} by 4=2\sqrt{4} = 2, and it does nothing at all to the center. Method A: mean still about 0.240, standard deviation about 0.0602=0.030\frac{0.060}{2} = 0.030. Method B: mean still about 0.385, standard deviation about 0.0692=0.0345\frac{0.069}{2} = 0.0345. Confirm both from the formula: 0.24(0.76)200=0.1824200=0.0009120.0302\sqrt{\frac{0.24(0.76)}{200}} = \sqrt{\frac{0.1824}{200}} = \sqrt{0.000912} \approx 0.0302 and 0.385(0.615)200=0.236775200=0.0011838750.0344\sqrt{\frac{0.385(0.615)}{200}} = \sqrt{\frac{0.236775}{200}} = \sqrt{0.001183875} \approx 0.0344.

  5. Note what part (d) demonstrates. After quadrupling nn, Method B's results cluster more tightly than ever around 0.385, so its bias has become easier to see and no smaller. A larger sample sharpens a wrong answer.

  6. (e) One sentence. Bias is the distance from the center of a method's sampling distribution to the parameter and depends only on the method, while sampling variability is the spread of that distribution around its own center and shrinks as nn grows.

(a) Method B, because its 200 results center at 0.385 while p=0.24p = 0.24; Method A centers at 0.241. (b) Bias 0.3850.240=0.145\approx 0.385 - 0.240 = 0.145 too high, from undercoverage: students living more than a mile away were left off the list and they almost never bike. (c) No. The mean of 200 results has standard deviation 0.0602000.0042\frac{0.060}{\sqrt{200}} \approx 0.0042, so a 0.001 miss is about a quarter of one standard deviation, while Method B misses by about 30. (d) A: mean about 0.240, standard deviation about 0.030. B: mean about 0.385, standard deviation about 0.0345. (e) Bias is how far the method's center sits from the parameter; sampling variability is the spread around that center, and only the spread responds to nn.

Problem 5

A school board wants the proportion of parents who support moving the high school start time from 7:20 a.m. to 8:15 a.m. A member drafts this question: "Given that teenagers who sleep longer perform better in school, do you support moving the high school start time to 8:15 a.m.?"

(a) Name the bias this wording creates and predict the direction. (b) Rewrite the question neutrally. (c) The board tests both wordings. It takes a simple random sample of 400 parents, splits them at random into two groups of 200, and reaches every parent by phone, with the caller introducing herself as calling on behalf of the school board and asking one group the draft question and the other group the rewrite. The draft gets 148 yes and the rewrite gets 106 yes. Find both proportions and the gap. (d) Why does splitting the sample at random matter for what part (c) proves? (e) The rewrite's result still is not the parameter. Name one bias that survives the fix and give its direction.

Show the worked solution
  1. (a) Name and direction. Selection was never the issue, so the lean is in the answers: this is response bias from question wording. The clause before the question hands the parent a reason to say yes and offers nothing on the other side, and it arrives while they are deciding. Support for the later start time will be overstated.

  2. (b) Rewrite it. Strip the argument, put both sides in the question, and offer a way out for parents with no view: "Do you support, oppose, or have no opinion about moving the high school start time from 7:20 a.m. to 8:15 a.m.?" The wording now names the change and nothing else, and neither answer is presented as the sensible one.

  3. (c) Compute both. Draft: p^=148200=0.74\hat{p} = \frac{148}{200} = 0.74, or 74%. Rewrite: p^=106200=0.53\hat{p} = \frac{106}{200} = 0.53, or 53%. The gap is 0.740.53=0.210.74 - 0.53 = 0.21, 21 percentage points, produced entirely by the sentence in front of the question.

  4. (d) Why the random split matters. The two groups of 200 came from one random sample and were formed by chance, so they are alike on average in every way except the wording each one heard. The wording is therefore the only systematic difference between them, and the 21 point gap can be attributed to it rather than to two different kinds of parent. Randomly assigning the wording turns this into an experiment about the question itself.

  5. (e) What survives. Response bias does, in a second form. The caller introduces herself as calling on behalf of the school board, which is the body proposing the change, and parents shade their answers toward whatever the caller seems to want. That pushes the reported support for the board's proposal upward, so 53% is best read as an upper estimate rather than as the parameter.

(a) Response bias from question wording, overstating support, because the sleep claim supplies a one-sided reason to say yes. (b) "Do you support, oppose, or have no opinion about moving the high school start time from 7:20 a.m. to 8:15 a.m.?" (c) 148200=0.74\frac{148}{200} = 0.74 and 106200=0.53\frac{106}{200} = 0.53, a gap of 21 percentage points. (d) The two groups were formed by chance from one random sample, so wording is the only systematic difference and the gap can be attributed to it. (e) Response bias in a second form: deference to a caller who says she represents the board, which pushes the reported support up.

Problem 6

A university health center mails a survey to a simple random sample of 500 students asking whether they got a flu shot this fall. 180 students answer and 126 of them say yes.

(a) Find the response rate and the proportion who said yes. (b) Name the bias, the mechanism, and the direction. (c) Without any assumption about the non-answerers, find the smallest and largest values the proportion vaccinated among the 500 selected students could take. (d) The health center reports the estimate with the usual 95% margin of error computed from its 180 answers, using the critical value 1.96. Find it and compare it to the range in part (c). (e) What should the health center do?

Show the worked solution
  1. (a) Two fractions. The response rate is 180500=0.36\frac{180}{500} = 0.36, or 36%, and among the answerers p^=126180=0.70\hat{p} = \frac{126}{180} = 0.70, or 70%.

  2. (b) Name it. The 500 were chosen by a random mechanism and 320 of them never answered, so this is nonresponse bias. A student who got the shot has a quick, creditable answer to give a health center that is plainly hoping to hear it, while a student who skipped it has a reason to leave the envelope unopened. The answerers lean vaccinated, so 0.70 overstates the proportion among the 500.

  3. (c) Bound it with no assumptions. 126 students definitely got the shot. If none of the 320 non-answerers did, the proportion among the 500 is 126500=0.252\frac{126}{500} = 0.252. If all 320 did, it is 126+320500=446500=0.892\frac{126 + 320}{500} = \frac{446}{500} = 0.892. So the parameter for these 500 lies somewhere in 0.252 to 0.892, a band 0.8920.252=0.6400.892 - 0.252 = 0.640 wide.

  4. (d) The reported margin of error. Using p^=0.70\hat{p} = 0.70 and n=180n = 180, the margin of error is 1.960.70(0.30)180=1.960.21180=1.960.0011667=1.96(0.03416)0.06691.96\sqrt{\frac{0.70(0.30)}{180}} = 1.96\sqrt{\frac{0.21}{180}} = 1.96\sqrt{0.0011667} = 1.96(0.03416) \approx 0.0669, about 6.7 percentage points. That would be reported as 70% plus or minus 6.7 points.

  5. Compare the two numbers. A 13.4 point interval sits inside a 64 point band of genuine uncertainty. The margin of error measures only the chance variation of a random sample of size 180; it assumes those 180 are a random sample from the population, which is exactly the assumption nonresponse breaks. No formula on the sheet widens an interval to cover bias, which is why a low response rate has to be reported alongside the estimate rather than buried.

  6. (e) The repair. Take a random subsample of the 320 non-answerers, say 60 of them, and pursue those 60 hard with phone calls or a visit until nearly all respond. That subsample represents the missing group, so combining it with the original 180 gives an estimate for all 500. Mailing the survey to 1,500 students instead would leave the response rate, and therefore the lean, exactly where it is. The margin of error arithmetic itself is Unit 3 work, and the margin of error calculator checks it.

(a) Response rate 180500=0.36\frac{180}{500} = 0.36; p^=126180=0.70\hat{p} = \frac{126}{180} = 0.70. (b) Nonresponse bias: vaccinated students have the easy, approved answer to send back to a health center, so 0.70 is too high. (c) From 126500=0.252\frac{126}{500} = 0.252 to 446500=0.892\frac{446}{500} = 0.892, a band 0.640 wide. (d) 1.960.211800.06691.96\sqrt{\frac{0.21}{180}} \approx 0.0669, about 6.7 percentage points, so a 13.4 point interval inside a 64 point band; the margin of error covers sampling variability only. (e) Chase a random subsample of the 320 non-answerers to near-complete response, rather than mailing more surveys.

Problem 7

A city council wants the proportion of residents who would use a proposed weekend shuttle to the beach. Staff set up a table at the downtown farmers market on three Saturday mornings and ask passersby, "Would you use a free weekend shuttle to the beach instead of fighting for parking?" 600 people stop and answer, and 462 say yes.

(a) This study has two distinct problems. Name both types and give the mechanism for each. (b) Give the direction each problem pushes the estimate. (c) The council later runs a proper survey: a simple random sample of 500 residents contacted by phone and mail with follow-ups, reaching 430 of them, of whom 129 say yes to a neutral version of the question. Find both proportions and the gap. (d) A staffer says the market result would have been trustworthy with 6,000 responses instead of 600. Compute what that change would actually do, and answer him.

Show the worked solution
  1. (a) Two separate stages, two separate problems. First, selection: nobody drew these 600 people. They were downtown on a Saturday morning and they chose to walk over to the table, so the sample is built from volunteers. That is voluntary response bias, layered on a frame that reaches only residents who visit a weekend farmers market. Second, the data: the question itself carries the words "free" and "instead of fighting for parking," which advertise the shuttle inside the question. That is response bias from question wording. The two are independent flaws, and repairing either one leaves the other in place.

  2. (b) Both push the same way. The people who stop at a table about a beach shuttle are the people who like the idea, and residents who go downtown on weekends are already the ones who travel for leisure, so the volunteer selection overstates support. The wording names a benefit and a pain point and mentions no cost or wait, so it also overstates support. The errors compound rather than cancel, and the estimate is inflated twice.

  3. (c) Compute both. Market: p^=462600=0.77\hat{p} = \frac{462}{600} = 0.77, or 77%. Proper survey: p^=129430=0.30\hat{p} = \frac{129}{430} = 0.30, or 30%, on a response rate of 430500=0.86\frac{430}{500} = 0.86. The gap is 0.770.30=0.470.77 - 0.30 = 0.47, 47 percentage points.

  4. (d) Price out the staffer's idea. Even pretending the 600 were a random draw from the market crowd, the standard deviation of p^\hat{p} is 0.77(0.23)600=0.1771600=0.000295170.0172\sqrt{\frac{0.77(0.23)}{600}} = \sqrt{\frac{0.1771}{600}} = \sqrt{0.00029517} \approx 0.0172. At 6,000 it is 0.17716000=0.0000295170.0054\sqrt{\frac{0.1771}{6000}} = \sqrt{0.000029517} \approx 0.0054.

  5. Say what that buys. Ten times the work tightens the spread by about 1.2 percentage points, around a number that is 47 percentage points away from the answer the random sample produced. Both flaws are properties of the table and the question, not of the count of people standing at it, so 6,000 market responses would be a more precise measurement of the wrong quantity.

(a) Voluntary response bias, since passersby chose themselves at a market that only some residents visit, and response bias from question wording, since "free" and "instead of fighting for parking" sell the shuttle inside the question. (b) Both overstate support, and they compound. (c) 462600=0.77\frac{462}{600} = 0.77 against 129430=0.30\frac{129}{430} = 0.30, a gap of 47 percentage points. (d) The standard deviation of p^\hat{p} would fall from about 0.0172 to about 0.0054, roughly 1.2 percentage points, while the 47 point lean is untouched; a bigger market sample measures the wrong quantity more precisely.

Problem 8

A fitness app publishes the claim "72% of Americans exercise at least three times a week." The evidence is an in-app pop-up asking users to report their weekly workouts. The app has 320,000 users, and 38,400 of them answered.

(a) Name the population the claim is about and the population this method can describe. (b) Identify three distinct sources of bias, with a mechanism and a direction for each. (c) Find the response rate. (d) The app's data scientist says the margin of error is under half a percentage point, so the estimate is reliable. Check the arithmetic with the critical value 1.96, then say precisely what that number does and does not cover. (e) Write one sentence the app could defensibly publish instead.

Show the worked solution
  1. (a) Two populations. The claim is about all Americans. The method can only describe people who had installed this fitness app and were still using it, which is a self-selected slice of the country, not a random sample of it.

  2. (b) Three mechanisms, all pushing up. Undercoverage, and severe: no American outside the app's user list could be selected, and people who install and keep a fitness app exercise far more than people who do not, so the frame itself sits above the national rate. Voluntary response: among users, answering the pop-up was optional, and a user proud of a consistent week taps through while a user who has not opened the app in a month dismisses it, so the answerers lean toward the more active users. Response bias: exercise is self-reported to an app whose whole purpose is encouraging exercise, and people round their own effort up. Every one of the three pushes the 72% higher than the truth, so no cancellation is available.

  3. (c) Response rate. 38400320000=0.12\frac{38400}{320000} = 0.12, or 12%. Nearly nine of every ten users said nothing.

  4. (d) Check the margin of error. With p^=0.72\hat{p} = 0.72 and n=38400n = 38400, 1.960.72(0.28)38400=1.960.201638400=1.960.00000525=1.96(0.002291)0.00451.96\sqrt{\frac{0.72(0.28)}{38400}} = 1.96\sqrt{\frac{0.2016}{38400}} = 1.96\sqrt{0.00000525} = 1.96(0.002291) \approx 0.0045, about 0.45 percentage points. The arithmetic is right, which is what makes the claim dangerous.

  5. Say what it covers. That 0.45 points is the spread you would see if this same procedure were repeated, and nothing more. It answers the question "how much would the answer bounce around from run to run," not "how far is the center from the truth." A method leaning 20 or 30 points high reports the same tiny margin of error every time it runs, because the formula contains p^\hat{p} and nn and no term at all for the quality of the sampling method. A large nn with a broken frame produces a confident number, not a correct one.

  6. (e) A defensible sentence. "Among the 38,400 users who chose to answer an in-app survey, 72% reported exercising at least three times a week." It names the group actually measured, flags that they chose to answer, and says "reported" rather than treating self-report as measured behavior. It makes no claim about Americans, which is the only claim the app wanted and the one its data cannot support. What a study does and does not license is worked through in can you generalize results.

(a) The claim is about all Americans; the method describes only active users of this one app. (b) Undercoverage (non-users cannot be selected, and app users exercise more), voluntary response (only the 12% who chose to tap through answered, and the active users are the proud ones), and response bias (self-reported exercise gets rounded up). All three push 72% above the truth. (c) 38400320000=0.12\frac{38400}{320000} = 0.12. (d) 1.960.2016384000.00451.96\sqrt{\frac{0.2016}{38400}} \approx 0.0045, about 0.45 percentage points, and it is correct; it measures run-to-run variability only and contains no term for a broken frame. (e) "Among the 38,400 users who chose to answer an in-app survey, 72% reported exercising at least three times a week."