AP Statistics Discussion Questions, by Topic
By Jude Wallis · Published
Below are forty-two discussion questions in six areas: study design and causation, sampling and bias, probability intuition, p-values, confidence intervals, and data ethics. Each group opens with the real tension underneath it, chosen so informed students can still land on opposite answers.
These questions sit mostly inside two of the four course practices: Collect Data, worth 20-30% of the exam, covers the study design and sampling groups, and Interpret Results, worth 25-35%, covers the p-value and confidence interval groups, since both practices ask a student to justify a conclusion rather than only compute one.
Why AP Statistics needs argument, not just arithmetic
Most AP Statistics practice trains one skill at a time: compute a standard deviation, carry out a test, interpret an interval. The exam's free response questions ask for something else. They hand you a scenario and ask you to justify a conclusion, which means weighing evidence and defending a judgment call, not retrieving a formula.
Discussion is the format that trains judgment directly, because a good discussion question has a correct underlying principle and a genuinely contestable application of it. A worksheet question has one right number. The forty-two questions below do not, or at least not one that every reasonable, well-informed student will reach the same way.
Use them as an opener before a unit, as a full period of review before a test, or as five-minute breaks inside a lecture when attention drops. For scenario-based practice with a single worked solution instead of a debate, see scope of inference practice and experimental design practice.
Study design and causation
Every AP Statistics student can recite the rule: random assignment supports a causal claim, random selection supports generalizing it to a population. The argument starts one level up, when a real study only gets one of the two, gets both imperfectly, or gets both cleanly but a decision cannot wait for a better one. Experiments vs observational studies and can you generalize these results cover the rule itself; these questions live in the gap the rule does not resolve for you.
- A large observational study finds that people who take a daily supplement live measurably longer. Should a doctor recommend the supplement on this evidence alone, or wait for a randomized trial? Name a gap in life expectancy that would be large enough for you to act on the observational finding anyway.
- A school randomly assigns half its incoming class to a study-skills seminar and the treated half's GPA rises more than the other half's. A parent whose child was assigned to the control group demands the seminar for their child too, since it "obviously works." Is refusing that request defensible?
- A drug cannot ethically be randomly assigned to teenagers to test whether it stunts growth, so all the evidence on the question is observational. Does the impossibility of ever running the clean experiment mean the observational conclusion stays provisional forever, or can enough converging observational studies substitute for the experiment nobody is allowed to run?
- In a randomized experiment, 12 of 60 people assigned to the treatment group drop out before it ends, and nobody drops out of the control group. Does the uneven dropout undo the causal conclusion, weaken it, or leave it intact, and where exactly is the line?
- A completely randomized, double-blind trial of a new sleep aid is run entirely on retired volunteers recruited from one clinic. Is the causal claim here strong or weak? Is the generalization strong or weak? Can a single study genuinely be excellent on one axis and close to worthless on the other?
- Is it possible for a well-designed experiment to be less useful to a policymaker than a well-designed observational study? Argue for a specific case where you would act on the correlational finding instead of waiting for a randomized one.
- Two teams study the same claim: one runs a randomized experiment on a small pool of volunteers, the other runs an observational study on records covering an entire population. Which result would you trust to set a policy meant for everyone, and why doesn't studying literally everyone settle the causal question by itself?
Sampling and bias
Bias is a property of a method, not of a person's intentions, which is why a well-meaning survey can be badly biased and a badly motivated one can, by luck, land close to the truth. The most common wrong turn in this unit is treating sample size as a stand-in for sample quality; does a bigger sample fix bias and how to identify the type of bias lay out why it is not, and these questions test whether that separation survives contact with a messier example.
- A city mails a satisfaction survey to a random sample of 5,000 households and gets 200 responses back. Is this a good estimate of citywide satisfaction, a biased one, or is there not yet enough information to say?
- A news app collects a genuinely random sample of its own users, all of whom opted in to political-news alerts, and reports the results as a poll of the public. What population, if any, does this result actually describe?
- Would you rather bet real money on a poorly randomized sample of 50,000 people or a carefully randomized sample of 500?
- A convenience sample of patients at one hospital happens, by coincidence, to match the county's age and income distribution exactly. Is the resulting estimate unbiased, or is that the wrong question to ask about a single sample?
- Every voluntary response sample this class has read about turned out to be biased toward the group with the strongest feelings. Can you design a voluntary response sample that avoids that particular bias, even in principle, or is it built into what "voluntary response" means?
- Splitting a sample by an irrelevant trait, like the first letter of a last name, and splitting it by a relevant one, like grade level, both technically produce a stratified random sample. Should the term require that the strata matter, or is that a separate claim from the sampling mechanics?
- Nonresponse is often assumed to bias a survey in one predictable direction, since people with strong feelings respond more. Can the direction of nonresponse bias be predicted before collecting the data, or only diagnosed after the fact? See selection bias and census vs sample survey for two of the mechanisms behind these questions.
Probability intuition
Correct probability answers routinely clash with strong gut instincts, and the disagreement in class is usually about which instinct deserves to be trusted, not about the arithmetic once it is laid out. Disjoint vs independent events and probability vs odds sort out the vocabulary; these questions test whether the vocabulary survives an intuition pump.
- Two independent events each have probability 0.5. Is knowing the outcome of the first one ever useful information about the second, even in a loose, non-mathematical sense, given that the formal answer is no?
- Two events are described as mutually exclusive. A student argues they must also be independent, since neither one "affects" the other. Is that a defensible use of the word independent, and exactly where does it break down?
- A forecast gives a 70% chance of rain, and it does not rain. Was the forecast wrong?
- In a class of 30 students, which is more surprising: that two people share a birthday, or that nobody does?
- A gambler who just watched a roulette wheel land on red six times in a row bets heavier on black next. Is there any version of that reasoning that holds up, even if the standard gambler's fallacy framing does not?
- A medical test is 99% accurate and comes back positive for a disease that only 1 in 10,000 people have. Is the patient more likely to have the disease or not, and why does the 99% figure feel like it should settle the question when it does not? Try probability of at least one as a warm-up before this one.
- Two events are independent in the statistical sense. Does that mean they are unrelated in every sense an ordinary English speaker would mean by "related"?
What a p-value does and does not say
The p-value is one of the most consistently misstated numbers in an introductory statistics course, and students who can compute one correctly still often reach for the wrong sentence the moment they are asked to explain it out loud. What does a p-value mean and p-value vs alpha give the correct wording; use these questions to see whether the class can hold onto it under pressure. A p-value is never the probability that the null hypothesis is true, and none of the correct answers below should be read that way.
- A p-value of 0.03 comes back for a new teaching method. A student says, "there's a 3% chance the method doesn't actually work." What exactly is wrong with that sentence, and can the class fix it in one sentence without losing anyone?
- Is a p-value of 0.049 meaningfully different from a p-value of 0.051, given that the standard cutoff treats them as opposite conclusions?
- Two studies test the same claim. One gets a p-value of 0.001 from a sample of 40; the other gets a p-value of 0.04 from a sample of 4,000. Which study gives stronger evidence of a real, useful effect?
- If a result is not statistically significant, has the study proven the null hypothesis true?
- A researcher runs 20 independent tests on the same dataset, and exactly one comes back significant at the 0.05 level. Should that one result be reported as a discovery?
- Does a smaller p-value always mean a bigger, more important effect?
- A significance level of 0.05 means roughly a 5% chance of a false alarm when the null hypothesis is actually true. Should every field use the same 0.05 cutoff, or should the number depend on what a false alarm costs? Bring in type 1 vs type 2 errors and how to choose a significance level before deciding.
Confidence intervals
A confidence interval reads like a statement about the one interval sitting in front of you, but it is actually a statement about the method that produced it, and this class has already met people who blur that line on purpose to sound more certain than the data supports. What 95% confidence means and confidence level vs confidence interval hold the precise wording.
- "There is a 95% probability that the true mean falls inside this specific interval." Once the interval has actually been calculated from the one sample in hand, is that sentence true?
- A pollster reports a candidate at 52% with a 3-point margin of error and calls it a clear lead. Is that a fair description of an interval running from 49% to 55%?
- Would a 99% confidence interval or a 90% confidence interval be more useful to a doctor deciding whether a new treatment's effect is real? Does higher confidence always mean more useful?
- Two different random samples from the same population produce two 95% confidence intervals that do not overlap at all. What, if anything, went wrong?
- If 100 different random samples from the same population each produced their own 95% confidence interval, roughly how many of the 100 intervals would actually contain the true parameter, and is that number a guarantee for any single one of the 100?
- A confidence interval for a difference between two groups contains zero. Does that mean the two groups are the same? Try two proportion confidence interval contains zero after the class settles this one.
- Making a confidence interval narrower always sounds like an improvement. Describe a way of narrowing one that would actually make the result worse, not better.
Data ethics
The ethical weight of a study is not obvious from its statistics alone, and two designs that are equally rigorous can differ sharply in whether either one should have been run at all, or run the way it was. Correlation vs causation and scope of inference supply the technical vocabulary that these questions push on from the ethical side.
- A randomized experiment could answer a valuable medical question, but only if a control group knowingly receives a treatment researchers already suspect does not work. Is a strong experimental design ever ethically identical to a pretext for withholding care?
- A company changes its app for millions of users without telling any of them, measures the effect, and reports the result as an experiment. Does the lack of consent change whether the statistics are valid, whether they are usable, or neither?
- A public health study wants a genuine random sample of an undocumented population to estimate need for services, which means finding and contacting people who have real reasons to avoid being found. Is refusing to attempt random sampling here the more ethical choice, even though it guarantees a weaker, less generalizable study?
- A dataset scraped from social media can answer a real research question, but nobody in it agreed to be studied. Does the usefulness of the answer change your view of whether it should have been collected?
- Is it more ethical to run an underpowered study that respects every constraint researchers face, or a well-powered one that cuts a corner on consent to get there?
- A company anonymizes its user data before selling it for research, but a separate public dataset makes part of it possible to re-identify. Whose responsibility is that: the company that anonymized it, the buyer, or the publisher of the second dataset?
- Misreading a p-value or a poll can cause real harm outside the classroom. When a statistic gets reported wrong, does the responsibility sit with the person who reported it, the person who misread it and acted on it, or both equally?
How to run the discussion
The format matters more than the question. A teacher who reads question 22 aloud and immediately calls on a volunteer gets one confident answer and silence from everyone who disagreed with it. A few habits change that.
Have students write a private answer before anyone speaks. Sixty seconds of silent writing forces a commitment before social pressure has a chance to flatten the room toward whoever spoke first. Then cold-call, and cold-call the quiet half of the room on purpose, since the students who write a different answer than the first speaker are usually the ones worth hearing from.
Withhold the correct framing until two genuinely different answers are on the table. A discussion that gets fixed after the first response was never a discussion. If the class reaches quick, confident agreement, that is the moment to introduce the counterexample you had ready, not to move on.
Require a restatement before a rebuttal. A student may not respond to another student's answer without first restating it in their own words, checked by the original speaker. This single rule kills most of the talking-past-each-other that makes a stats discussion feel unproductive.
Timebox each question to somewhere between four and eight minutes, including the ones with a clean correct answer, like most of the p-value group. Six of these in a fifty-minute period, with a short break for a written answer between them, uses the period well without turning into a lecture with occasional pauses.
Do not grade for the right answer. Grading the content of an opinion on data ethics, or on how large an observational gap has to be before you act on it, guarantees that students report the position that seems safest rather than the one they hold. If a grade is needed, grade for whether a student restated an opposing view accurately, which is checkable and does not reward picking the popular side.
Close each question with the sentence the class would defend in writing, not a verdict. For the questions with a real correct principle, such as the p-value and confidence interval groups, that sentence should match the standard wording; students can then rehearse writing it as a full justification on experimental design practice, p-value interpretation practice, or confidence interval practice. For the ethics questions, the closing sentence is often "reasonable people land in different places, for these specific reasons," and that is a complete, honest answer.
Worked walkthrough: running the p-value sentence question
A student says a p-value of 0.03 means there is a 3% chance the new teaching method does not actually work. Run this as a five-minute discussion that gets the class to a correct sentence without the teacher simply announcing the fix.
Give the class sixty seconds to write, individually and silently, whether they agree with the student's sentence and one reason why. This forces a private commitment before social pressure sets in.
Cold-call two students who wrote opposite answers and have each read their sentence aloud without interruption from anyone else.
Ask the class what would have to be true for the 3% to be a probability about the teaching method rather than about the data. It would mean the method already had a known probability of working before the study ran, which nothing in a significance test supplies.
Redirect to what the p-value actually assumes: it is calculated by assuming the method has no effect, then asking how unusual a result at least this large would be under that assumption.
Have the class build the corrected sentence together, word by word, comparing it against the student's original phrase to name exactly which noun phrase moved: chance about the data given the assumption, never chance about the assumption given the data.
Close by asking whether 0.03 alone tells you the method is worth adopting in every classroom, to open the door into the significance-versus-importance questions later in the same group.
The corrected sentence: if the teaching method truly has no effect, there is roughly a 3% chance that a sample like this one would show a difference at least this large from chance alone. That 3% attaches to the data given the assumption of no effect, never to the assumption given the data. See what does a p-value mean for the general form of this correction.
Worked walkthrough: running the confidence interval overlap question
Two different random samples from the same population produce two 95% confidence intervals that do not overlap at all. Run this as a discussion that reaches the real mechanics of coverage instead of a shrug.
Draw both intervals on the board as two number lines and ask for a show of hands: did someone make a mistake?
In pairs, have students generate every explanation they can in ninety seconds, such as a calculation error, an unusually unlucky sample, a population that changed between the two samples, or ordinary sampling variability, and share one per pair without judging any of them yet.
Reframe with the definition itself: a 95% confidence interval means that if this same method were used on many different random samples, about 95% of the resulting intervals would contain the true parameter, not that intervals from different samples should agree with each other.
Run the number: with a 95% method, roughly one interval in twenty misses the true parameter by chance alone even when nothing else has gone wrong, so two non-overlapping intervals from two honestly random samples is an uncommon but entirely possible outcome.
Push the harder case: ask what evidence, beyond the non-overlap itself, would make the class suspect the population genuinely changed between the two samples rather than chance alone. This is usually where the class splits, and that split is the actual point of the question.
Close by naming what would count as a real error instead: mismatched sample sizes reported as equal, two different formulas used for the two intervals, or an arithmetic mistake, since those are checkable and chance alone is not.
Non-overlapping 95% intervals from two honest random samples of the same population are not automatic evidence of an error, since about one interval in twenty misses the true parameter by chance alone under repeated random sampling. The discussion should end with the class able to name what additional evidence, beyond the non-overlap, would be needed to conclude something other than chance happened. See why is my confidence interval so wide and confidence level vs confidence interval for the mechanics behind it.
Frequently asked questions
Do I need a prepared correct answer for every question?
For the probability, p-value, and confidence interval groups, yes: those have a real correct principle underneath the disagreement, and the discussion should end there. For most of the data ethics group and several of the study design questions, the honest close is that informed people land in different places for specific, statable reasons, and that is a complete answer rather than a dodge.
How long should each question run?
Four to eight minutes including the silent-write step, even for questions with a clean correct answer. Longer than that and the class either resolves it and stalls, or the same two or three students end up carrying the whole exchange.
Should these be graded?
Not for the position a student takes. Grading the content of an opinion pushes students toward the answer that seems safest to say out loud rather than the one they actually hold, which defeats the purpose. If a grade is needed at all, grade whether a student restated an opposing view accurately before responding to it.
What if the class agrees immediately and there is no real argument?
Quick, confident agreement is the cue to introduce the counterexample rather than move to the next question, for example the gambler who has a defensible reason buried inside an indefensible instinct, or the honest pair of non-overlapping confidence intervals. Have that counterexample ready before you ask the question.
Can this replace free response practice?
No, but it prepares for it. A discussion trains the reasoning; the free response section grades the written justification. Pair each group with matching practice, such as experimental design practice or p-value interpretation practice, and have students write the closing sentence from the discussion as a full justification afterward.