Scope of inference practice problems (8 solved)

By Jude Wallis · Updated

This set has 8 problems on what a study lets you conclude and about whom. Random assignment is what supports a cause-and-effect claim; random selection is what supports generalizing to a population. Every answer states both halves in the phrasing a free-response answer needs.

AP Statistics: Unit 1 (topics 1.10 The Investigative Question Revisited and Data Collection, 1.13 Experimental Design). These problems cover the justification objectives in Unit 1 of the Fall 2026 AP Statistics course: topic 1.10 asks you to justify the appropriateness of generalizations for a statistical study, and topic 1.13 asks you to justify the appropriateness of conclusions based on a well-designed experiment. Unit 1 is 20 to 30% of the multiple-choice section.

What these problems build

Two separate randomizations answer two separate questions, and most of the points lost on these items come from letting one answer the other. Random assignment of treatments to units is what supports a cause-and-effect conclusion, because it makes the treatment groups alike on average with respect to every other variable, which is what reduces the potential for confounding. Random selection of units from a population is what supports generalizing the result to that population. Neither one does the other's job.

Read a study by asking those two questions in order, then place it in the grid.

Random assignment usedNo random assignment
Units randomly selectedcause and effect, generalizes to the population sampledassociation only, generalizes to the population sampled
Units not randomly selectedcause and effect, for these units onlyassociation only, for these units only

The bottom row is not a failure, it is a smaller claim. When units are not randomly selected, the conclusion still applies, just to individuals similar to the ones in the study. When treatments are not assigned, a confounding variable can always sit behind the pattern, and naming one is part of a full answer: to qualify, the variable has to be associated with the explanatory variable and also affect the response, which is what makes its effect impossible to separate from the effect under study.

Every solution here writes the conclusion in two labeled halves, which is the shape a free-response answer needs.

  1. Causation. Because the units were randomly assigned to the treatments, a difference this large is evidence that the explanatory variable caused the change in the response. Or: because no treatment was assigned, this study shows an association only.
  2. Generalization. Because the units were randomly selected from a stated population, the conclusion applies to that population. Or: because the units volunteered, the conclusion applies only to individuals similar to them.

For the underlying distinctions, read experiments vs observational studies, correlation vs causation, and can you generalize results. The course topics are 1.10 The Investigative Question Revisited and Data Collection and 1.13 Experimental Design. To build the designs themselves rather than judge them, work experimental design practice.

Problem 1

For each study, say whether a cause-and-effect conclusion is supported and name the population the result covers.

(a) A cardiology group keeps a registry of 18,000 patients with one type of arrhythmia. It randomly selects 300 of them, then randomly assigns 150 to a new dosing schedule and 150 to the standard schedule.

(b) A polling firm randomly selects 1,000 people from a state's registered voter file and records how many hours each one volunteers per month and whether each one owns a dog.

(c) A gym owner posts a sign-up sheet, gets 40 volunteers, and randomly assigns 20 to each of two stretching routines before a timed sprint.

(d) A newsletter writer asks her subscribers to reply with their daily coffee intake and their sleep quality. 900 replies come back.

Show the worked solution
  1. (a) Both randomizations are present. Treatments were randomly assigned, so cause and effect is supported, and the 300 patients were randomly selected from the 18,000-patient registry, so the conclusion generalizes to that registry. This is the top-left cell, the strongest position a study can be in.

  2. (b) Selection is random, assignment is absent. Nobody assigned dog ownership; the firm recorded what already existed, so this is an observational study and only an association can be reported. The 1,000 were randomly selected from the voter file, so that association generalizes to the people on that file. Note the population is the registered voters on the file, not all adults in the state, since unregistered adults could not be selected.

  3. (c) Assignment is random, selection is not. The routines were randomly assigned, so a difference in sprint times can be attributed to the routine. The 40 people volunteered from one gym's sign-up sheet, so the conclusion covers those 40 and people like them, not gym members in general.

  4. (d) Neither randomization is present. Coffee intake was not assigned, so association only. The 900 chose to reply to an open invitation, so the result describes people similar to the subscribers who chose to reply, and even within her subscriber list the volunteers may not resemble the quiet majority.

  5. Line up the four. Reading the grid, (a) is cause and effect plus generalization, (b) is generalization without cause, (c) is cause without generalization, and (d) is neither. The first question is always what was assigned; the second is always who was selected.

(a) Cause and effect supported; generalizes to the 18,000 patients in the registry. (b) Association only; generalizes to the registered voters on the state file. (c) Cause and effect supported; applies only to these 40 volunteers and people like them. (d) Association only; applies only to people similar to the subscribers who chose to reply.

Problem 2

An auto insurer offers a phone app that tracks braking and speed, and customers choose whether to install it. Over one year, 1,820 of the 40,000 customers who installed the app filed a claim, compared with 3,900 of the 60,000 customers who did not.

(a) Find each claim rate and express the app group's rate as a percentage of the other group's. (b) Is this an experiment or an observational study? Say how you can tell. (c) The insurer's press release says the app makes drivers safer. Name a confounding variable and explain both links that make it confounding. (d) State the conclusion the insurer can defend, in the two-part form.

Show the worked solution
  1. (a) Two rates. App users: 182040000=0.0455\frac{1820}{40000} = 0.0455, or 4.55%. Non-users: 390060000=0.065\frac{3900}{60000} = 0.065, or 6.50%. The ratio is 0.04550.065=0.70\frac{0.0455}{0.065} = 0.70, so the app group's claim rate is 70% of the other group's, a 30% lower rate and a difference of 0.0650.0455=0.01950.065 - 0.0455 = 0.0195, or 1.95 percentage points.

  2. (b) Study type. Customers decided for themselves whether to install the app, so no treatment was imposed and the insurer recorded what happened. That makes it an observational study, and an observational study can establish an association but not a cause.

  3. (c) Name a confounder. Driving caution, meaning the habits a driver already had before hearing about the app. First link: cautious drivers are far more willing to let an insurer watch their braking and speed, because they expect the recording to flatter them, so caution is associated with installing the app. Second link: cautious drivers crash less regardless of any app, so caution affects the claim rate directly. Because the careful drivers piled into the app group, the effect of the app and the effect of habits those drivers already had cannot be separated, and the 1.95 point gap could belong to either one.

  4. Check what does not qualify. The lower premium the insurer gives app users is a consequence of installing the app, not something sitting behind the choice, so it is part of how the app might work rather than a rival explanation. A confounding variable has to be tied to the explanatory variable and also affect the response on its own.

  5. (d) The defensible conclusion. Causation: because customers chose the app rather than being randomly assigned to it, this study shows only that app use and claim rate are associated, and the association may be explained by the kind of driver who opts in. Generalization: the 100,000 customers are this insurer's whole book of business rather than a random sample of drivers, so the association describes this insurer's customers and should not be extended to drivers generally. To defend the causal claim the insurer would have to randomly assign customers to install the app or not.

(a) 182040000=0.0455\frac{1820}{40000} = 0.0455 against 390060000=0.065\frac{3900}{60000} = 0.065; the app group's rate is 0.04550.065=0.70\frac{0.0455}{0.065} = 0.70, or 70% of the other rate, 1.95 percentage points lower. (b) Observational, because customers chose whether to install the app and no treatment was imposed. (c) Driving caution: cautious drivers are more willing to be monitored, and cautious drivers crash less anyway, so the app's effect and the drivers' pre-existing habits cannot be separated. (d) App use is associated with a lower claim rate among this insurer's customers; no causal claim is supported without random assignment, and no extension to drivers generally is supported without random selection.

Problem 3

A sleep researcher recruits 90 volunteers from a university's psychology participant pool and randomly assigns 45 to a screen curfew, no phone or laptop after 9 p.m., and 45 to their usual evening habits, for two weeks. Mean time to fall asleep is 12.4 minutes in the curfew group and 21.8 minutes in the usual-habits group, a difference far larger than the study's chance variation.

(a) Find the difference in means and state which group fell asleep faster. (b) Write the causal half of the conclusion. (c) Write the generalization half of the conclusion. (d) The researcher wants to headline the result as "a screen curfew helps adults fall asleep faster." What single change to the study would justify the word adults, and why is that change hard to make? (e) Why does random assignment support a causal claim here even though nobody was randomly selected?

Show the worked solution
  1. (a) The difference. 21.812.4=9.421.8 - 12.4 = 9.4 minutes, so the curfew group fell asleep 9.4 minutes faster on average.

  2. (b) Causation. Because the 90 volunteers were randomly assigned to the two conditions, the two groups were alike on average with respect to age, caffeine habits, workload, and every other variable at the start. A 9.4 minute difference larger than chance variation would produce is therefore evidence that the screen curfew caused the faster sleep onset.

  3. (c) Generalization. The 90 were volunteers from one university's psychology participant pool, not a random sample of any population, so the conclusion applies to people similar to these volunteers: university students who sign up for sleep studies. It does not extend to adults in general, whose evening screen use, work schedules, and ages differ from a student pool's.

  4. (d) What would justify "adults." Only random selection of the participants from the population of adults would support that word, since generalization comes from how the units entered the study, not from how treatments were assigned. That change is hard because an experiment requires people to follow a two-week protocol, and a randomly selected adult can refuse. Anyone who agrees has volunteered, which puts the study back where it started. The realistic middle path is to recruit volunteers from several settings, say a hospital, a workplace, and a retirement community, so that the group at least resembles a wider range of adults, and to say plainly that the sample was not random.

  5. (e) Why the two questions stay separate. Random assignment works inside the group you already have: it distributes the extraneous variables carried by those 90 people evenly across the two conditions, so the only systematic difference between the groups is the curfew. That argument never mentions the rest of the world, so it holds no matter how the 90 arrived. Random selection is what connects the 90 to a larger population, and it is missing, so the causal finding is real but local.

(a) 21.812.4=9.421.8 - 12.4 = 9.4 minutes faster for the curfew group. (b) Because the volunteers were randomly assigned, the 9.4 minute difference is evidence that the screen curfew caused faster sleep onset. (c) Because they were volunteers from one psychology participant pool rather than a random sample, the conclusion applies only to people similar to them. (d) Randomly selecting participants from the adult population, which is hard because selected adults can decline a two-week protocol and everyone who agrees is again a volunteer. (e) Random assignment balances the extraneous variables carried by these 90 people across the two groups, an argument that never refers to anyone outside the study.

Problem 4

A state agriculture department draws a simple random sample of 900 of the state's 12,000 registered beekeepers and records, for each one, whether the hives sit within a mile of a sunflower field and whether the colony survived the winter. Of the 240 keepers with hives near sunflower fields, 204 colonies survived. Of the other 660, 429 survived.

(a) Find the survival rate in each group, the overall survival rate, and the gap. (b) Which cell of the selection-by-assignment grid is this study in? (c) State what may and may not be concluded, and about whom. (d) Name a confounding variable and explain both links. (e) Sketch a study that would support the causal claim, and say what population it would cover.

Show the worked solution
  1. (a) Three rates. Near sunflowers: 204240=0.85\frac{204}{240} = 0.85, or 85%. Not near: 429660=0.65\frac{429}{660} = 0.65, or 65%. Overall: 204+429900=6339000.7033\frac{204 + 429}{900} = \frac{633}{900} \approx 0.7033, about 70.3%. The gap is 0.850.65=0.200.85 - 0.65 = 0.20, 20 percentage points.

  2. (b) Locate it in the grid. The 900 were randomly selected from the 12,000 registered beekeepers, so random selection is present. Nobody assigned hive locations; keepers had already put their hives where they wanted, so random assignment is absent. This is the top-right cell: generalization yes, causation no.

  3. (c) The two halves. Generalization: because the 900 were a simple random sample of the state's 12,000 registered beekeepers, the 20 point difference generalizes to that population of registered beekeepers. Causation: because hive placement was not assigned, the study shows an association between sunflower proximity and winter survival, not that proximity causes survival.

  4. (d) A confounder. The keeper's experience and management practice. First link: keepers who run larger, more deliberate operations are the ones who negotiate placements next to sunflower fields, while a hobbyist puts two hives in the back garden, so experience is associated with proximity. Second link: experienced keepers treat for mites on schedule, feed in autumn, and insulate for winter, all of which raise survival regardless of what is growing nearby. The two effects arrive together, so neither the sunflowers nor the management can be credited with the 20 points.

  5. (e) A design that would earn the causal claim. Take the same random sample of registered keepers, ask each to supply two comparable hives, then randomly assign one hive of each pair to a sunflower-adjacent site the department arranges and the other to a site with no sunflowers nearby. Random assignment inside each pair balances the keeper's own management across the two conditions, because the same keeper tends both hives. With random selection and random assignment both present, the study would support a causal conclusion for the state's registered beekeepers. The catch is that participation is voluntary, so keepers who decline would drag it back toward a volunteer sample.

(a) 204240=0.85\frac{204}{240} = 0.85 and 429660=0.65\frac{429}{660} = 0.65, overall 6339000.703\frac{633}{900} \approx 0.703, a 20 percentage point gap. (b) Random selection, no random assignment: the top-right cell. (c) The association generalizes to the state's 12,000 registered beekeepers, but no causal claim is supported because placement was not assigned. (d) Keeper experience and management: experienced keepers both secure sunflower-adjacent sites and manage mites and feeding well, so their effect on survival cannot be separated from the sunflowers'. (e) Have each sampled keeper supply two comparable hives and randomly assign one of each pair to a sunflower-adjacent site; that adds random assignment to the existing random selection and would support causation for the state's registered beekeepers.

Problem 5

A state department of education randomly selects 60 of the state's 480 middle schools, then randomly assigns 30 of them to a 90-minute daily reading block and 30 to their current schedule, for one school year. Mean school reading score gain is 6.2 points under the new block and 2.9 points under the current schedule, a difference larger than chance variation would produce.

(a) What fraction of the state's middle schools were selected, and how many are in each group? (b) Identify the experimental units and explain why that matters for the wording of the conclusion. (c) Write the full two-part conclusion. (d) A principal in a neighboring state reads the study and adopts the block. Is that supported? Explain. (e) A reporter writes that the reading block raises any individual student's score by 3.3 points. Correct the sentence.

Show the worked solution
  1. (a) Counts. 60480=0.125\frac{60}{480} = 0.125, so 12.5% of the state's middle schools were selected, and the 60 split into 602=30\frac{60}{2} = 30 schools per group.

  2. (b) The units are schools. A schedule was assigned to a school, not to a student, so the experimental unit is the school and the response is the school's mean gain. That controls the wording: the study compares 30 schools with 30 schools, and every conclusion is a statement about schools adopting a schedule. Writing it as a statement about students would treat thousands of students as separately assigned units, which they were not.

  3. (c) The two halves. Causation: because the 60 schools were randomly assigned to the two schedules, the two groups of schools were alike on average with respect to size, funding, and prior performance, so a 6.22.9=3.36.2 - 2.9 = 3.3 point difference in mean gain larger than chance variation is evidence that the reading block caused the higher gain. Generalization: because the 60 schools were randomly selected from the state's 480 middle schools, that causal conclusion extends to the state's middle schools.

  4. (d) The neighboring state. Not supported by this study. Random selection was from one state's 480 middle schools, so the population the result covers is that state's middle schools. A neighboring state may differ in curriculum, school day length, and student population, and nothing in this design speaks to it. The principal is making a judgment that his schools resemble the ones studied, which may be sensible but is not a statistical conclusion. Replicate in his own state to support it.

  5. (e) Fix the reporter's sentence. Two things are wrong: the unit and the certainty. The 3.3 points is a difference between two group means of schools, not a promise for any one student or even any one school, and school means vary around it. Rewrite: "Among the middle schools in the study, those randomly assigned the 90-minute reading block gained 3.3 more points on average than those keeping the current schedule." That names the units, keeps the claim on means, and keeps the random assignment visible.

(a) 60480=0.125\frac{60}{480} = 0.125, or 12.5%, with 30 schools per group. (b) The units are schools, since the schedule was assigned to a school, so every conclusion is a statement about schools and their mean gains. (c) Because the schools were randomly assigned, the 3.3 point difference in mean gain is evidence the reading block caused the higher gain; because they were randomly selected from the state's 480 middle schools, that conclusion extends to the state's middle schools. (d) No. The population sampled was one state's middle schools, so the result does not reach the neighboring state without replication there. (e) "Among the middle schools in the study, those randomly assigned the 90-minute reading block gained 3.3 more points on average than those keeping the current schedule."

Problem 6

A researcher takes a simple random sample of 500 students from one large university's list of 21,000 enrolled students. Of the sampled students, 185 report studying in a group at least weekly and have a mean GPA of 3.28; the other 315 have a mean GPA of 3.09. The campus paper runs the headline "Group study raises grades for college students nationwide."

(a) Find the difference in mean GPA and verify that the overall sample mean GPA is 3.16. (b) The headline overreaches in two separate ways. Name each one and say which feature of the study it violates. (c) Rewrite the headline so that it claims exactly what the study supports. (d) Name a confounding variable and explain both links. (e) What one change would fix the first overreach, and what one change would fix the second?

Show the worked solution
  1. (a) Two calculations. The difference is 3.283.09=0.193.28 - 3.09 = 0.19 GPA points. The overall mean is the weighted average 185(3.28)+315(3.09)500=606.8+973.35500=1580.15500=3.1603\frac{185(3.28) + 315(3.09)}{500} = \frac{606.8 + 973.35}{500} = \frac{1580.15}{500} = 3.1603, which rounds to 3.16 as reported.

  2. (b) Two separate overreaches. The word "raises" claims cause. Nobody assigned students to study in groups; the researcher recorded a choice students had already made, so this is an observational study and only an association is supported. The word "nationwide" claims a population. Random selection was from one university's list of 21,000 students, so the sample represents that university and nothing beyond it. One flaw is about assignment, the other about selection, and fixing one leaves the other untouched.

  3. (c) A headline that fits. "Students at this university who study in groups have higher GPAs on average." It reports an association, names the population that was actually sampled, and drops any suggestion that group study produced the difference. Adding the size is fair: the gap was 0.19 GPA points.

  4. (d) A confounder. Course load difficulty, or more precisely the major a student is in. First link: group study is normal in engineering and pre-med sequences with weekly problem sets and rare in majors built on solo reading and essays, so major is associated with group study. Second link: majors differ in how strictly they grade, so major affects GPA on its own. Prior academic preparation works the same way: students who arrive with stronger study habits both seek out study groups and earn higher grades. Either variable would produce a GPA gap with no contribution from group study at all.

  5. (e) Two different repairs. To earn the word "raises," randomly assign the sampled students to a group study requirement or to solo study for a term and compare GPAs, since assignment is what balances major and preparation across the groups. To earn the word "nationwide," draw the sample randomly from a national list of enrolled college students rather than from one university's roster. A study with both changes would support the original headline; a study with either one alone supports only half of it.

(a) 3.283.09=0.193.28 - 3.09 = 0.19 GPA points, and 185(3.28)+315(3.09)500=1580.15500=3.16033.16\frac{185(3.28) + 315(3.09)}{500} = \frac{1580.15}{500} = 3.1603 \approx 3.16. (b) "Raises" claims cause without random assignment, since group study was self-selected; "nationwide" claims a population beyond the one university whose roster was sampled. (c) "Students at this university who study in groups have higher GPAs on average, by 0.19 points." (d) Major or prior preparation: both are tied to whether a student joins study groups and both affect GPA directly, so their effect cannot be separated from group study's. (e) Randomly assign group versus solo study to fix the causal overreach; randomly select from a national roster of enrolled students to fix the population overreach.

Problem 7

A clinic recruits 200 patients with knee osteoarthritis from its own waiting list and randomly assigns 100 to a 12-week supervised exercise program and 100 to usual care. During the 12 weeks, 34 patients in the exercise group quit and 6 in the usual-care group quit. The clinic analyzes only the patients who finished: mean 6-minute walk distance is 462 meters for the exercise finishers and 418 meters for the usual-care finishers.

(a) How many patients are in each analyzed group, and what are the two dropout rates? (b) What would the study have supported if all 200 had been measured? (c) Explain why comparing the finishers breaks the causal claim. (d) Predict the direction of the resulting error in the 44 meter difference. (e) State what the clinic can defensibly report, and what it should do next time.

Show the worked solution
  1. (a) Counts and rates. Analyzed: 10034=66100 - 34 = 66 exercise patients and 1006=94100 - 6 = 94 usual-care patients. Dropout rates: 34100=0.34\frac{34}{100} = 0.34, or 34%, against 6100=0.06\frac{6}{100} = 0.06, or 6%, a difference of 28 percentage points. The reported gap is 462418=44462 - 418 = 44 meters.

  2. (b) With everyone measured. Treatments were randomly assigned, so a difference larger than chance variation would support the conclusion that the exercise program caused the longer walk distance. Generalization would still be limited, because the 200 came from one clinic's waiting list rather than a random sample, so the conclusion would cover patients similar to these.

  3. (c) Why the finishers are not the randomized groups. Random assignment created two groups of 100 that were alike on average. The groups being compared are no longer those: they are 66 patients and 94 patients who were filtered by an event that happened after assignment and depended on the treatment received. Quitting a supervised exercise program is not a coin flip; it happens to the patients whose knees hurt most and who are least able to keep up. The two analyzed groups were therefore formed partly by the patients themselves, which is exactly what random assignment existed to prevent, so the comparison has become observational.

  4. (d) The direction. The 34 who quit the exercise arm are, on average, the ones doing worst on it, and removing them leaves the fitter, more tolerant two thirds. Almost nobody was filtered out of the usual-care arm by comparison, since only 6 left. The exercise group's mean is therefore lifted by the removal in a way the usual-care group's is not, so 44 meters overstates the program's effect. Note also that a 34% dropout rate is itself a result: a program most patients cannot finish is worth less than one they can.

  5. (e) What to report and what to change. Defensibly: among patients who completed 12 weeks, the exercise finishers walked 44 meters farther on average than the usual-care finishers, an association that may reflect who was able to finish rather than the program itself, and the exercise program had a 34% dropout rate against 6% for usual care. Next time, measure the 6-minute walk on all 200 patients at week 12 whether or not they stuck with the program, and analyze them in the groups they were assigned to, which keeps the randomization intact and keeps the causal conclusion available.

(a) 66 and 94 analyzed; dropout rates 34100=0.34\frac{34}{100} = 0.34 and 6100=0.06\frac{6}{100} = 0.06, 28 percentage points apart. (b) That the program caused the longer walk distance, for patients similar to this clinic's waiting-list patients, since the 200 were not randomly selected. (c) The 66 and the 94 were not the groups chance created; patients filtered themselves out after assignment, in a way that depended on the treatment, so the comparison is now observational. (d) The 44 meters overstates the effect, because the patients doing worst on the program are the ones who left it while almost none left usual care. (e) Report an association among completers plus the 34% versus 6% dropout rates, and next time measure all 200 at week 12 and analyze by assigned group.

Problem 8

A city added protected bike lanes downtown in one year. In the year before, there were 84 cyclist collisions across 2.1 million bike trips. In the year after, there were 63 collisions across 2.8 million trips. The city also lowered the downtown speed limit from 30 to 25 miles per hour in the same year.

(a) Find the collision rate per 100,000 trips for each year and the percent decrease. (b) Why is comparing the raw counts 84 and 63 misleading here? (c) Classify the evidence and say what it supports. (d) Explain why the speed limit change is a confounding variable, using both links. (e) Describe a study that would support the causal claim, and name the population it would cover.

Show the worked solution
  1. (a) Rates first. Before: 842100000=0.00004\frac{84}{2100000} = 0.00004 per trip, which is 0.00004×100000=4.000.00004 \times 100000 = 4.00 collisions per 100,000 trips. After: 632800000=0.0000225\frac{63}{2800000} = 0.0000225 per trip, which is 2.252.25 collisions per 100,000 trips. The decrease is 4.002.25=1.754.00 - 2.25 = 1.75 per 100,000 trips, and 1.754.00=0.4375\frac{1.75}{4.00} = 0.4375, a 43.75% decrease in the rate.

  2. (b) Why counts mislead. The denominator changed. Trips rose from 2.1 million to 2.8 million, an increase of 2.82.12.10.333\frac{2.8 - 2.1}{2.1} \approx 0.333, or about 33%, so more riding was happening in the second year. A count of 63 against 84 understates the improvement, because those 63 collisions are spread over a third more cycling. The rate per trip is the comparable quantity.

  3. (c) Classify the evidence. No treatment was randomly assigned to anything: the city changed its streets and watched. No random selection took place either, since this is one city's downtown observed at two points in time. This is an observational before-and-after study, in the bottom-right cell, so it supports an association between the year and the collision rate for this city's downtown and nothing more.

  4. (d) The speed limit as a confounder. First link: it changed at exactly the same time as the bike lanes, so it is perfectly associated with the explanatory variable, which here is before versus after. Second link: lower vehicle speeds give drivers more stopping time and reduce collisions on their own. Because the two changes are locked together, no amount of data from these two years can say which one produced the drop, or how much each contributed. A city-wide downward trend in collisions, or better weather in the second year, would work the same way.

  5. (e) A design that would earn the causal claim. Suppose the city has 260 comparable downtown corridors. Randomly select 40 of them, randomly assign 20 to receive protected bike lanes and 20 to keep their current layout, hold the speed limit at 25 miles per hour on all 40, and compare collisions per 100,000 trips after one year. Random assignment makes the two sets of corridors alike on average in traffic volume, width, and weather, so a difference in rate can be attributed to the lanes; random selection from the 260 corridors extends that conclusion to the city's comparable downtown corridors. Note that the experimental units are corridors, so the conclusion is about corridors, and the study covers this one city rather than cities in general.

(a) 842100000×100000=4.00\frac{84}{2100000} \times 100000 = 4.00 and 632800000×100000=2.25\frac{63}{2800000} \times 100000 = 2.25 collisions per 100,000 trips, a decrease of 1.75, or 1.754.00=43.75%\frac{1.75}{4.00} = 43.75\%. (b) Trips rose about 33%, so the raw counts compare different amounts of cycling and understate the improvement. (c) An observational before-and-after study with neither random assignment nor random selection, so it supports an association for this city's downtown only. (d) The speed limit dropped at the same time as the lanes went in, so it is tied to the before-versus-after variable, and lower speeds cut collisions on their own; the two effects cannot be separated. (e) Randomly select 40 of the city's 260 comparable corridors, randomly assign 20 to get lanes and 20 to keep the current layout with the speed limit held fixed, and compare rates after a year; that supports causation for this city's comparable downtown corridors.