Investigative question practice (8 problems)

By Jude Wallis · Published

This set has 8 problems on writing an investigative question and picking the data collection that matches it. A question is answerable when it names a measurable variable, a parameter with a direction or a range, and the population the conclusion will cover.

AP Statistics: Unit 1 (topics 1.1 Introducing Statistics: What Can We Learn from Data?, 1.10 The Investigative Question Revisited and Data Collection, 1.13 Experimental Design). These problems cover topic 1.10 of the Fall 2026 AP Statistics course, which splits a well-formed investigative question into components that guide data collection, the analysis, and the conclusion, and which names the census, experiment, observational study, and survey. Topic 1.1 is where the question is first posed and fixed before the data arrive, and the design write-ups in problems 5 and 8 draw on topic 1.13. Unit 1 is 20 to 30% of the multiple-choice section, and Question 1 of the free-response section is Multi-Focus on Practices 1 and 2, Formulate Questions and Collect Data.

What these problems build

An investigative question is answerable when a stranger could take it, collect data, and know when they had answered it. Most wonderings do not clear that bar, and the same three moves fix them every time.

The course splits a well-formed investigative question into three components.

  1. The collection component, phrased in terms of the variable or variables of interest. It tells a data collector what to write down and about whom.
  2. The analysis component. For a hypothesis test it names the parameter and the direction of the alternative: not equal, greater than, less than, association, or not independent. For a confidence interval it names the parameter and the goal of estimating it within a range.
  3. The conclusion component. It says what kind of conclusion is coming and which population it applies to, and, for an experiment that uses random assignment, whether a cause-and-effect conclusion is allowed.

Run those past a vague sentence and the gap shows up immediately. "Do our students get enough sleep?" fails the first component, because enough is a judgment and not something you can record for a student. "Is the bike share working?" fails all three, because working is not a variable, no parameter is named, and no population is fixed.

The payoff is that the answerable version usually names its own collection method. A question about a condition somebody can impose points at an experiment. A question about what people report points at a survey. A question about a variable nobody can assign points at an observational study. Problems 1, 5, and 8 run the wondering-to-question move; problem 2 picks the collection; problem 4 writes the analysis component; and problems 3, 6, and 7 catch questions the data on hand cannot answer.

This is Practice 1: Formulate Questions, the smallest of the four course practices at 5 to 10%, and Question 1 of the free-response section is Multi-Focus on Practices 1 and 2, so it shares that question with Collect Data. The topics are 1.1 Introducing Statistics and 1.10 The Investigative Question Revisited and Data Collection.

Matching a question to a way of collecting data

The course names four ways to collect data, and the question decides which one you get.

Way to collectThe signal in the situationWhat it can answer
Censusevery item or individual in the population can be measuredthe parameter itself, with nothing left to infer
Surveythe data are what people report to a standard set of questionsa parameter of the population sampled, plus associations among what was asked
Observational studythe explanatory variable cannot or should not be assignedassociation only, for the population sampled
Experimentthe researcher can impose the condition on the unitscause and effect, for the units in the study, when the treatments are randomly assigned

A survey is one kind of observational study, so a survey never earns a causal claim either. The course also splits observational studies by direction in time: a prospective study follows units forward from the moment they enter, and a retrospective study reads records of things that already happened.

Only an experiment imposes treatments, and that single fact carries most of the classification questions. Two follow-ups separate the rest: can you measure everyone (census) or only some (sample), and do the data come from people answering questions (survey) or from records and measurements (not a survey). The wider contrast is in experiments vs observational studies and experiment vs survey.

What the question is allowed to promise

The third component is where these problems overlap the scope-of-inference work, and it runs on two separate randomizations.

Random selection of units from a population is what lets the conclusion reach that population. Random assignment of treatments to units is what lets the conclusion name a cause. A question whose conclusion component promises either one has to be backed by the matching randomization in the collection plan, or the question is not answerable as written.

That is why a wondering can be turned into two perfectly good questions with completely different answers, as problem 8 does with the same bike share. Write the conclusion component first if it helps: knowing you want to say something causal about all 240 stations tells you immediately that you need both randomizations, and therefore an experiment run on a random sample of stations.

For the conclusion wording itself, work can you generalize results and scope of inference practice. For the designs the causal questions call for, work experimental design practice, and for the sampling the descriptive ones call for, sampling methods practice.

Frequently asked questions

What makes an investigative question answerable?

It names a variable someone could actually record, it names the parameter along with either a direction (not equal, greater than, less than, association, not independent) or the goal of estimating within a range, and it names the population the conclusion will apply to. A sentence missing any of the three leaves a data collector guessing.

Is a survey an observational study?

Yes. A survey collects data from people using a standard set of questions and imposes no treatment, which makes it one kind of observational study. That is why no survey supports a causal claim, whatever its size. Survey is the more specific label when the data come from people answering questions.

How do I tell when a question calls for an experiment?

Two things have to be true at once: the question is causal, and somebody could impose the condition on the units. If the explanatory variable is one people already have, such as years of night-shift work or where a family lives, no experiment is available and the honest question asks about association instead.

Can I change the investigative question after seeing the data?

No. The question fixes the variable, the parameter, and the direction before collection, and a result cannot be evidence for a question that was chosen because of that result. Problem 6 prices it: run 30 comparisons at a 5% significance level with nothing there and you expect 30×0.05=1.530 \times 0.05 = 1.5 of them to look significant.

Where does this appear on the exam?

Formulating questions is Practice 1, banded at 5 to 10%, the smallest of the four course practices. Question 1 of the free-response section is Multi-Focus on Practices 1 and 2, so it shares that question with Collect Data. The content sits in topic 1.10, at the end of Unit 1, which is 20 to 30% of the multiple-choice section.

Problem 1

A counselor at a high school with 1,480 students says, "I want to know whether our students get enough sleep." The district wellness policy names 8 hours on a school night as its target. The counselor can pull 120 students out of class for one period, and she plans to select those 120 at random from the school's enrollment list.

(a) Explain why her sentence is not yet an answerable investigative question, and say which part of it is already fine. (b) Write the collection component, the part phrased in terms of the variable of interest. (c) Write the analysis component as a confidence interval question, and name the parameter. (d) Write the analysis component as a hypothesis test question, and name the parameter and the direction. (e) Write the conclusion component, and find what fraction of the school is in the sample.

Show the worked solution
  1. (a) Hold the sentence against the three components. The variable is missing. "Enough sleep" is a judgment, not something a data collector can record: two students could disagree about whether 7 hours is enough and both be answering honestly, so the sentence does not say what to write down. What is already fine is the population. "Our students" is a defined group of 1,480 people with an enrollment list behind it, which is more than most wonderings supply.

  2. (b) The collection component. Replace the judgment with a variable someone can measure: "On a typical school night, how many hours of sleep do students at this school get?" Hours of sleep on a school night is quantitative, every student has a value, and a data collector now knows exactly what to ask. Any wording that names hours of sleep on a school night and the students at this school does the job.

  3. (c) The analysis component, interval version. A confidence interval question names the parameter and says the goal is to estimate it within a range: "What is the mean number of hours of sleep on a school night for the 1,480 students at this school, and what range of values is plausible for it?" The parameter is μ\mu ("mu"), the mean for all 1,480. It is not the mean of the 120 who get asked, which is xˉ\bar{x} ("x-bar"), a statistic.

  4. (d) The analysis component, test version. A test question names the parameter and the direction of the alternative. The 8 hour target comes from the stem, so a direction is available: "Is the mean number of hours of sleep on a school night for students at this school less than 8 hours?" Parameter μ\mu again, direction less than. The interval version and the test version are different questions about the same parameter, and the counselor has to choose one before collecting rather than after seeing the numbers.

  5. (e) The conclusion component and the fraction. The conclusion will describe the mean sleep of the 1,480 students at this school, and it will be a description rather than a cause, because nobody is assigning anyone a bedtime. It reaches all 1,480 because the 120 are randomly selected from the enrollment list. The fraction sampled is 12014800.081\frac{120}{1480} \approx 0.081, about 8.1% of the school.

(a) It names no measurable variable, since "enough sleep" is a judgment rather than something you can record; the population, the 1,480 students at this school, is already well defined. (b) "On a typical school night, how many hours of sleep do students at this school get?" (c) "What is the mean number of hours of sleep on a school night for the 1,480 students at this school, and what range of values is plausible for it?" The parameter is μ\mu, the population mean. (d) "Is the mean number of hours of sleep on a school night for students at this school less than 8 hours?" Parameter μ\mu, direction less than. (e) The conclusion describes the mean sleep of all 1,480 students with no causal claim, and it reaches them because the 120 were randomly selected; 12014800.081\frac{120}{1480} \approx 0.081, about 8.1% of the school.

Problem 2

For each situation, name the way of collecting data it calls for, choosing from census, survey, observational study, and experiment, and give the one feature of the situation that forces your choice. Use the most specific label that fits, since a survey is itself one kind of observational study.

(a) A district needs to know exactly how many of its 62 school buses fail this month's brake inspection. (b) A cafeteria manager wants to know whether opening a second, clearly labeled serving line shortens the mean time students wait for lunch. She can open the second line on some days and leave it closed on others, and she decides which days. (c) A student council has one homeroom period, enough to reach about 200 of the school's 900 students, and wants to estimate the proportion who would use a late bus. (d) A researcher wants to compare rates of a sleep disorder between adults who worked night shifts for at least five years and adults who never did. Nobody can be put on night shifts for five years for a study. (e) A safety team registers 900 newly licensed teen drivers, records which of them chose to take an optional advanced driving course, and follows all 900 forward for three years, counting crashes.

(f) Situations (d) and (e) are both observational studies, one prospective and one retrospective. Say which is which and how you can tell.

Show the worked solution
  1. (a) Census. The population is 62 buses sitting in the district's own yard, and the district needs the exact number rather than an estimate. A census records information from every item in the population, so nothing is inferred and there is no sampling variability to account for. It is not a survey, because buses do not answer questions.

  2. (b) Experiment. The manager can impose the condition: she decides which days the second line opens. Imposing a treatment is the whole definition of an experiment, so an experiment is available to her. Deciding the days by hand is not enough to earn the causal claim, though. If she opens the second line on the days she expects to be busy, the treatment moves with the crowd, and the explanation that busy days are simply the days she opens it survives untouched. What removes that explanation is random assignment: draw which days get the second line with a random mechanism and every other variable, the day of the week, the menu, the weather, the size of the lunch period, is spread across the two groups by chance rather than by her expectations. A shorter mean wait is then left with only two explanations, the second line or the luck of the draw, and that is what a causal claim rests on. Note where the treatment lands: the waits are timed on students, but the condition is assigned to a day, and the experimental unit is whatever the treatment is assigned to, so the units here are days.

  3. (c) Survey. The data are what people report about what they would do, which no record contains, and the council can reach only 2009000.222\frac{200}{900} \approx 0.222, about 22.2%, of the school, so it is a sample rather than a census. Whether that 22.2% supports a statement about all 900 depends on how the 200 are chosen, which the stem does not say. Students who happen to sit in one homeroom are a convenience sample, and the estimate would then describe those 200 only.

  4. (d) Observational study. The explanatory variable is years of night-shift work, and the stem says outright that it cannot be assigned. With no assignment there is no experiment, so the researcher compares people who already differ, and the comparison can report an association only.

  5. (e) Observational study. The drivers chose the course for themselves; the team registered them and watched, but assigned nothing. Following units forward in time does not make a study an experiment. The only thing that would is imposing the course on a randomly chosen half of the 900.

  6. (f) Direction in time. Situation (d) reads work history that already happened, so it is retrospective. Situation (e) starts at registration and follows the same 900 drivers forward for three years, so it is prospective. Both are observational, and neither earns a causal claim, which is the point: the direction in time changes the logistics, not the scope of the conclusion.

  7. Read the set as a decision tree. First ask whether anyone imposes a condition, which picks out (b) alone. Then ask whether everyone can be measured, which picks out (a). Then ask whether the data are people's answers to a standard set of questions, which picks out (c). What is left is observational, and the only remaining question is which way it faces in time.

(a) Census: the 62 buses are the whole population and the district needs the exact count. (b) Experiment: the manager imposes the condition by deciding which days the second line opens, and the units are days; to earn a causal claim she has to assign those days at random rather than by hand, since otherwise busy days can be the days she chooses. (c) Survey: the data are what people report, and only about 200 of 900 can be reached, 2009000.222\frac{200}{900} \approx 0.222, about 22.2%; whether it reaches all 900 depends on how those 200 are chosen, which the stem does not say. (d) Observational study: night-shift history cannot be assigned. (e) Observational study: the drivers chose the course themselves. (f) (d) is retrospective, reading work history that already happened; (e) is prospective, following the 900 drivers forward for three years.

Problem 3

A city news site with about 210,000 daily readers publishes an article about a proposed transit levy and puts a poll box at the bottom: "Do you support the levy? Yes / No." Over one day, 4,812 readers answer and 3,051 of them click Yes. The next morning the site's headline reads "City backs transit levy, 63% say yes."

(a) Find the proportion who clicked Yes and the share of the daily readership that answered at all. (b) Which of these three investigative questions can this data set answer? (i) What proportion of the readers who answered the poll support the levy? (ii) What proportion of the city's residents support the levy? (iii) Did reading the article change readers' support for the levy? (c) Name the kind of sample this poll produced, and say what it does to the estimate. (d) The editor says the poll would be trustworthy if 50,000 readers had answered instead of 4,812. Respond. (e) Write an investigative question about the city that the site could actually answer, and say what it would have to collect.

Show the worked solution
  1. (a) Two proportions. Of those who answered, 305148120.634\frac{3051}{4812} \approx 0.634, so 63.4% clicked Yes, which is where the headline's 63% comes from. The people who answered are 48122100000.023\frac{4812}{210000} \approx 0.023 of the daily readership, about 2.3%, so roughly 2100004812205,000210000 - 4812 \approx 205{,}000 readers who could have clicked said nothing.

  2. (b) Sort the three questions. (i) Yes, and exactly. The 4,812 who answered are not a sample of themselves, they are all of themselves, so 63.4% answers (i) with no inference and no margin of error attached. (ii) No. The respondents were not randomly selected from the city's residents; they selected themselves by clicking, and they had to be reading this site's transit article to be in the pool at all. Nothing in the data connects them to the city. (iii) No. Nobody was assigned to read or not read the article, and no support was measured beforehand, so there is neither a comparison group nor a before value.

  3. (c) Name the sample. It is a voluntary response sample: every respondent decided for themselves to be counted. People who feel strongly about a levy click far more often than people who feel mildly, so 63.4% is pushed off the readership's true share in a consistent direction. That is bias, a systematic error, not the random wobble a random sample would have.

  4. (d) The editor's fix. It does not work. A larger voluntary response poll measures the same wrong quantity more precisely: the tilt lives in who chooses to click, and that is unchanged at 50,000. Sample size shrinks variability and does nothing to bias, which is the argument in does a bigger sample fix bias. What the site would have on 50,000 responses is a confident wrong number.

  5. (e) A question the site could answer. "What proportion of the city's registered voters support the transit levy, and what range of values is plausible for that proportion?" To answer it, draw a random sample from the city's voter registration list, contact those people directly, read every one of them the same wording, and record how many could not be reached. The parameter is pp, the proportion of registered voters who support the levy. Note where the conclusion stops: registered voters, because that is the list the sample came from, and not all residents, since anyone not on the list had no chance of selection.

(a) 305148120.634\frac{3051}{4812} \approx 0.634, 63.4% of respondents, and 48122100000.023\frac{4812}{210000} \approx 0.023, about 2.3% of the daily readership. (b) Only (i), and exactly, because the 4,812 are all of the respondents rather than a sample of them. (ii) fails because they selected themselves and are not connected to the city, and (iii) fails because nothing was assigned and nothing was measured beforehand. (c) A voluntary response sample, which biases the estimate in a consistent direction because people with strong feelings click more often. (d) No. A bigger voluntary response poll estimates the same wrong quantity more precisely; size reduces variability, not bias. (e) "What proportion of the city's registered voters support the transit levy, and what range of values is plausible for it?" Collect it by contacting a random sample drawn from the voter registration list with identical wording, and stop the conclusion at registered voters.

Problem 4

Each investigative question below states a goal. For each, say whether the analysis component points to a confidence interval or a hypothesis test, name the parameter the question is about, and, where a test is called for, name the direction of the alternative.

(a) A quality manager wants a range of plausible values for the mean weight, in grams, of the almonds her machine puts in a bag. (b) A shipping contract promises that at most 5% of packages arrive late. The operations lead wants to know whether his late rate is above that. (c) A researcher wants to know whether the proportion of adults who compost differs between two counties. (d) A principal wants to know whether favorite sport, with three categories, is associated with class year, with four categories, among the students at her school.

(e) One of the four cannot be answered with an interval estimate in this course. Say which, and why.

Show the worked solution
  1. (a) Interval, no direction. The manager asked for a range of plausible values, which is the confidence interval form of the analysis component: name the parameter, and say the goal is to estimate it within a range. The parameter is μ\mu ("mu"), the mean weight in grams of all the bags the machine fills. Nothing is being decided against a stated value, so there is no direction to give.

  2. (b) Test, direction greater than. The contract supplies the value to test against, 0.05, and the words above that supply the direction. The parameter is pp, the proportion of all this operation's packages that arrive late, and the alternative is p>0.05p > 0.05. Keep the parameter a proportion of packages: a count of late packages and a percentage of one particular sample are both something else.

  3. (c) Test, direction not equal. Differs is two-sided. The researcher has not said which county she expects to be higher, so the alternative points both ways. The parameter is p1p2p_1 - p_2, the difference between the two counties' population proportions, and the alternative is that this difference is not 0. Two populations are in play, so the question is not answerable until it names both counties.

  4. (d) Test, direction association or not independent. That is the fourth kind of direction the analysis component allows, and it is the one that does not come with a single number. The claim is that favorite sport and class year are related among this school's students, so the alternative is that the two variables are associated, or equivalently not independent.

  5. (e) The one with no interval. It is (d). An interval estimate needs a parameter to place an interval around, and with 3 sports crossed against 4 class years there is no single number that captures the relationship. Contrast (c): a difference between two proportions is a number, so it could be estimated with an interval instead of tested, and there the choice comes from the stated goal rather than from the data. That is worth noticing, because it means the analysis component is a decision you make, not one the data makes for you.

(a) Confidence interval; parameter μ\mu, the mean bag weight in grams; no direction. (b) Hypothesis test; parameter pp, the proportion of packages that arrive late; direction greater than, with the alternative p>0.05p > 0.05. (c) Hypothesis test; parameter p1p2p_1 - p_2, the difference between the two counties' proportions; direction not equal. (d) Hypothesis test; the claim is that favorite sport and class year are associated, or not independent, and there is no single numeric parameter. (e) (d), because 3 sports crossed with 4 class years leaves no single number to put an interval around, while (c)'s difference between two proportions is a number and could be estimated instead of tested.

Problem 5

A coach says, "I want to know whether our new warm-up works." His squad is 48 players. He has last season's warm-up and this season's new one, and he can decide which players do which.

(a) Name two different measurable responses the word works could mean here, and say why the study cannot leave both open. (b) He settles on 20-meter sprint time in seconds, recorded right after the warm-up. Name the explanatory variable and its levels. (c) Which way of collecting data does his question call for, and what one feature of the situation makes it available? (d) Name the experimental units and give the group sizes for an even split. (e) The athletic director wants to roll the new warm-up out to all 14 schools in the district. Explain why this study does not answer his question, and name the change that would.

Show the worked solution
  1. (a) Two responses, and why you have to choose. Works could mean a faster 20-meter sprint time in seconds, or fewer hamstring injuries over a season, or a longer sit-and-reach distance in centimeters. Any two of those count. They are different studies: a sprint time is one afternoon's measurement taken on every player, while injuries are rare enough that a season and far more than 48 players would be needed before a difference meant anything. The reason both cannot stay open is that the investigative question has to be fixed before the data arrive. A coach who measures five things and reports the one that came out well has not answered a question, he has chosen one, and problem 6 puts a number on how often that produces a result from nothing.

  2. (b) The explanatory variable. It is the warm-up routine, a categorical variable with two levels: the new routine and last season's routine. Those two levels are the treatments. Sprint time is the response, measured after the treatment, and no second factor is in play, so this is one explanatory variable with two levels rather than a crossed design.

  3. (c) The collection method. An experiment. The feature that makes it available is stated in the stem: he can decide which players do which warm-up. Imposing the condition is what separates an experiment from everything else, so the causal question is open to him. Earning the answer takes the random assignment set up in part (d): letting players pick a routine, or handing the new one to the players he thinks need it, would leave the difference in sprint times explained just as well by which players ended up in which group.

  4. (d) Units and sizes. A warm-up is assigned to a player, so the experimental units are the 48 players. An even split is 482=24\frac{48}{2} = 24 players per routine. Do the assignment with a mechanism rather than by preference: number the players 1 to 48, generate random integers from 1 to 48 ignoring repeats until 24 distinct numbers appear, give those 24 the new routine, and give the remaining 4824=2448 - 24 = 24 last season's.

  5. (e) The director's question. This study covers 48 players on one squad, and they were not randomly selected from anything. Random assignment earns the causal claim for those 48 and players like them; it says nothing about the district's other 13 schools, whose players differ in age, sport, and training history. The change that would answer him is random selection: draw players, or whole squads, at random from across the 14 schools, then randomly assign the two warm-ups within that sample. Random assignment and random selection do two different jobs, and he is asking for the one this study does not have.

(a) Any two measurable responses, for example 20-meter sprint time in seconds and hamstring injuries per season. Both cannot stay open because the question must be fixed before the data arrive, and they need different studies: a sprint time is one afternoon's measurement on every player, while injuries need a season and far more players. (b) Explanatory variable: warm-up routine, with two levels, the new routine and last season's, which are the treatments. (c) An experiment, available because the coach can decide which players do which warm-up. (d) The 48 players are the units, 482=24\frac{48}{2} = 24 per routine, assigned by random number generator rather than by preference. (e) The 48 are one squad and were not randomly selected, so the causal finding covers them and players like them; random selection of players or squads from across the 14 schools is what would extend it to the district.

Problem 6

A researcher is handed a data file on the 1,200 students at one high school. It records each student's GPA and whether they belong to each of 30 school clubs. He had no question in mind when the file arrived. He compares mean GPA between members and non-members for each club in turn, 30 comparisons, using a 5% significance level each time. Two come out significant. He writes up one of them as "Students in debate club have higher GPAs" and presents it as the question the study set out to answer.

(a) Suppose none of the 30 clubs is really related to GPA. About how many of the 30 comparisons would you expect to come out significant anyway? (b) Which requirement of an investigative question did he break? (c) His defense is that the finding is real because the data are real and the arithmetic is correct. Answer him. (d) Even if debate club membership and GPA really are associated in this file, name the second thing his sentence claims that the data cannot support. (e) Say what he should do with the debate club finding.

Show the worked solution
  1. (a) Count the false alarms you should expect. If no club is really related to GPA, each comparison still has a 5% chance of coming out significant on its own, so across 30 of them you expect 30×0.05=1.530 \times 0.05 = 1.5 significant results from chance alone. That figure needs nothing beyond the 5% per comparison; it does not require the 30 comparisons to be unrelated to each other, which they are not, since a student can belong to several clubs. He found 2. That is about what running 30 comparisons produces when there is nothing there.

  2. (b) The requirement he broke. An investigative question has to be posed before the data are collected and analyzed, and it must not be changed once the results are visible. He had no question at all, so the data chose one for him. The written-up question exists because of its result, which is exactly what stops that result from being evidence for it. The first component of an investigative question names the variable of interest in advance, and here the variable of interest was picked last.

  3. (c) Answer the defense on its own terms. Every one of the 30 comparisons uses real data and correct arithmetic, and none of that is in dispute. The problem is which comparison got reported. He ran 30 and published the ones that looked good, and part (a) says roughly 1.5 would look good even if every club were unrelated to GPA. Correct arithmetic on a comparison selected because of how it turned out tells you the arithmetic was right, not that the comparison was worth making.

  4. (d) The second unsupported claim. The population. His sentence says "students," which reads as students in general, while the file covers the 1,200 students at one high school and nobody was randomly selected from anywhere. Even a solid association here reaches those 1,200 and students like them, and the sentence has to say so.

  5. (e) What to do with it. Treat it as a question rather than as an answer. Write the investigative question down first, before any new data: "Is mean GPA higher for debate club members than for non-members at this school?" That fixes the variable, the parameter, and the direction in advance. Then collect a fresh set of data and answer that one question. A pattern spotted in a file is a reason to run a study, not the result of one. Note that the fresh study is still observational, since nobody assigns students to debate club, so even a confirmed difference stays an association.

(a) 30×0.05=1.530 \times 0.05 = 1.5 comparisons, so about 1 or 2, which is what he got. (b) The investigative question must be fixed before the data are analyzed, and he had none: the data chose the question, so the result that chose it cannot also be evidence for it. (c) Real data and correct arithmetic are not the issue; the issue is that he ran 30 comparisons and reported the ones that came out well, and about 1.5 of 30 come out well by chance alone. (d) The population. The file is one high school's 1,200 students with no random selection, so the claim reaches those students and students like them, not students in general. (e) Write it down as an investigative question in advance, naming the variable, parameter, and direction, then collect fresh data to answer that one question; and note that the answer will still be an association, since debate club membership is never assigned.

Problem 7

A high school's records already contain, for every one of this year's 1,140 students: class year, whether the student rides a school bus, days absent, and final GPA. Of the 1,140, 486 ride a bus. For each proposed investigative question, say whether the existing records answer it, and if not, say exactly what would have to be collected.

(a) What is the mean number of days absent for this year's students at this school? (b) Are days absent and GPA associated at this school? (c) Does riding the bus cause students to miss more school? (d) How many hours of sleep do this school's students get on a school night? (e) Do students at this school miss more school than students statewide?

(f) Find the proportion of the school that rides a bus, and say whether that proportion is a parameter or a statistic.

Show the worked solution
  1. (a) Answerable, and exactly. Days absent is recorded for every one of the 1,140 students, so the records are a census of the school on that variable. Compute the mean and you have the parameter itself. There is no sample, no interval, and no margin of error, because nobody was left out and nothing is being inferred.

  2. (b) Answerable as an association. Both variables sit in the records for all 1,140 students, so the relationship between them can be described for this school with no new collection at all. What the records cannot supply is a direction of influence, because neither variable was assigned to anyone.

  3. (c) Not answerable. Nobody assigned bus riding: students ride or do not ride depending on where they live, and where a family lives is tied to plenty of things that affect attendance on their own. Earning the word cause would take an experiment that randomly assigns students to ride or not, which a school cannot run. The version of (c) these records can answer is (b)'s shape: are absences associated with riding the bus at this school.

  4. (d) Not answerable. Sleep is not in the file, and no analysis creates a variable nobody recorded. To answer it, collect the variable: ask a random sample of the school's students how many hours they slept on a stated school night, which supports a conclusion about all 1,140, or ask every student for a census. Either way this is new collection, and since the school holds no record of anyone's sleep it has to come from the students themselves, which in practice means a survey.

  5. (e) Not answerable. The records are a census of this school and contain no student from any other school, so there is nothing to compare against. You would need absence data for a random sample of the state's students, or a published statewide figure, and the comparison would then be between this school and that population rather than between this school and itself.

  6. (f) The proportion and its label. 48611400.426\frac{486}{1140} \approx 0.426, about 42.6% of the school rides a bus, and the other 1140486=6541140 - 486 = 654 do not. Because the records cover every student at the school, that number describes the whole population of interest, so it is a parameter for this school, not a statistic. The same arithmetic would produce a statistic if the 1,140 were a sample of something larger. Parameter or statistic depends on what the number describes, never on how it was computed.

(a) Yes, exactly: days absent is recorded for all 1,140, so the mean is a parameter with nothing to infer. (b) Yes, as an association, since both variables are in the records; the records cannot say which one moves the other. (c) No. Bus riding was never assigned, so only an experiment could support cause, and the records support an association at most. (d) No. Sleep is not a recorded variable, so it has to be collected, by surveying a random sample of the school's students about a stated school night. (e) No. The records hold no student from any other school, so a statewide comparison needs absence data from a random sample of the state's students or a published statewide figure. (f) 48611400.426\frac{486}{1140} \approx 0.426, about 42.6%, and it is a parameter, because the records cover every student at the school.

Problem 8

A city transportation office says, "We need to figure out whether the bike share is working." Two facts are on hand: the system has 12,000 registered riders on a list the office maintains, and 240 docking stations that are comparable in size and location type. The office can afford to contact 400 riders, and it can afford to build covered shelters at 45 stations.

(a) Turn the wondering into an answerable descriptive question, giving all three components, and name the way of collecting data it calls for. (b) Turn the same wondering into an answerable causal question about shelters, giving all three components, and name the way of collecting data it calls for. (c) For (b), describe the randomization and the group sizes, using 90 of the 240 stations. (d) State what each of the two studies licenses the office to conclude, and about whom. (e) Why can neither study, on its own, tell the office whether the bike share is working?

Show the worked solution
  1. (a) The descriptive question. Collection component: in the past 30 days, how many bike share trips did a registered rider take for which a car was available to them at the time and they chose the bike instead? That is quantitative, and it is something only the rider can report, which is why the office's own trip records cannot supply it. Note what the wording avoids: asking how many trips a rider "would otherwise have driven" asks for a guess about a trip that never happened, and two riders making the same trip could answer it differently and both be honest, which is the objection problem 1 used to reject "enough sleep." Whether a car was available is something the rider knows. Analysis component: what is the mean number of such car-available bike trips for the 12,000 registered riders, and what range of values is plausible for it? The parameter is μ\mu, that population mean. Conclusion component: the conclusion describes the 12,000 registered riders and claims no cause, and it reaches all 12,000 only if the 400 are randomly selected from the office's list. The method is a survey of 400120000.033\frac{400}{12000} \approx 0.033, about 3.3% of the riders.

  2. (b) The causal question. Collection component: how many trips start at a docking station over an 8-week period, and does that station have a covered shelter? Analysis component: is the mean number of trips started per station over 8 weeks greater for stations with a covered shelter than for stations without one? The parameter is the difference between the two treatment means, and the direction is greater than. Conclusion component: because shelters will be randomly assigned, a difference can be attributed to the shelter, and because the stations in the study will be randomly selected from the 240, that conclusion reaches the city's 240 comparable stations. The method is an experiment, available because the office can decide which stations get shelters.

  3. (c) The randomization, in two stages. First select. Number the 240 comparable stations 1 to 240, use a random number generator to draw integers from 1 to 240 ignoring repeats until 90 distinct numbers appear, and those 90 stations are the study's units. That is 90240=0.375\frac{90}{240} = 0.375, 37.5% of the comparable stations. Then assign. Number those 90 from 1 to 90, draw distinct integers from 1 to 90 ignoring repeats until 45 appear, build covered shelters at those 45, and leave the remaining 9045=4590 - 45 = 45 as they are. Count trips started at each of the 90 over the same 8 weeks and compare the two mean trip counts. The experimental units are stations, because a shelter is built at a station, so every conclusion is a statement about stations.

  4. (d) What each licenses. The survey: because the 400 riders were randomly selected from the list of 12,000, the estimate of the mean number of car-available bike trips generalizes to all 12,000 registered riders, and because nothing was assigned, it describes them rather than explaining anything. The experiment: because shelters were randomly assigned, a higher mean trip count at sheltered stations is evidence the shelter caused it, and because the 90 stations were randomly selected from the 240, that causal conclusion reaches the city's 240 comparable stations. The experiment has both randomizations, which is the strongest position a study can be in and the rarest.

  5. (e) Why working stays unanswered. Working is not a variable. Each study answers one narrow question about one measurable thing, and the office still has to decide in advance what it will count as working: car-available bike trips, trips per station, rider retention, cost per trip. That decision belongs to the first component of the investigative question, and it belongs before any data arrive, because a measure chosen after the results are in is chosen for its result. What the two studies do show is that the choice of measure decides the collection method: a condition the office can impose points at an experiment, and something only a rider can report points at a survey.

(a) Collection: how many bike share trips in the past 30 days a registered rider took with a car available to them at the time, which is something a rider knows, unlike a guess about what they would otherwise have done. Analysis: the mean number of such trips for the 12,000 registered riders, estimated within a range, so the parameter is μ\mu. Conclusion: a description of all 12,000, no cause, valid only if the 400 are randomly selected. Method: a survey of 400120000.033\frac{400}{12000} \approx 0.033, about 3.3% of the list. (b) Collection: trips started per station over 8 weeks, and whether the station has a shelter. Analysis: is the mean trips per station greater for sheltered stations, so the parameter is the difference in treatment means and the direction is greater than. Conclusion: causal, and reaching the city's 240 comparable stations. Method: an experiment. (c) Draw 90 of the 240 at random ignoring repeats, 90240=0.375\frac{90}{240} = 0.375, then draw 45 of those 90 at random for shelters, leaving 9045=4590 - 45 = 45 unchanged; compare mean trips started per station over the same 8 weeks, with stations as the units. (d) The survey generalizes a description to all 12,000 registered riders; the experiment supports a causal claim and generalizes it to the 240 comparable stations, because it has both randomizations. (e) Working is not a variable, so the office has to name the measure it will use before collecting, and that choice is what selects the method.