AP Statistics Review Games Worth Class Time
By Jude Wallis · Published
Ten review games built to force writing or reasoning rather than recall, from a State-Plan-Do-Conclude relay to two games run on the site's own draggable interactives. Each lists timing, prep, what it trains, and its best-fit units, and the final section sequences them into one review day.
These are review-day activities for the whole course rather than content tied to one topic. The relay and the condition audit apply the inference structure taught across Unit 3 and Unit 4; the error hunt and free response jigsaw work on any unit's material; and kill the correlation is specific to the regression content in Unit 5.
Why these games and not another round of multiple choice
A practice set tells you whether a student can recognize a correct answer among four options. Half of the AP Statistics exam does not ask that. Section II asks a student to write a full State-Plan-Do-Conclude paragraph, name a condition and defend it, and interpret a number in the context of the actual problem, and no amount of multiple choice drilling builds the habit of writing that paragraph under time pressure.
The ten games below are sorted by what they force a student to do rather than by unit, because the same format works across several units. Two are built directly on this site's interactives, so "reveal the answer" is not a teacher's whiteboard but a live simulation the class already trusts. Timing, team size, and prep are listed for each so a teacher can drop one into a fifteen minute block without redesigning it.
1. The State-Plan-Do-Conclude relay
Time: 10 minutes per round. Teams: four, one whiteboard and one marker each. Prep: a bank of inference scenarios with the numbers already given (sample size, mean or count, standard deviation, claimed parameter).
Call a scenario. Student one writes State: the parameter in context and the hypotheses. Student two writes Plan: the correct procedure and checks every condition by name. Student three writes Do: the standardized test statistic and the p-value, or the interval. Student four writes Conclude: the decision compared against alpha, or the interval's capture claim, answered in the context of the problem. Only a complete board scores, so the fourth student's sentence counts as much as the arithmetic.
Trains: the exact four-part structure the FRQ guide asks for, with each part assigned to a different student so the relay exposes which stage a team is weak at. Fits: Unit 3 and Unit 4, where every inference free response follows this shape.
2. The error hunt
Time: 12 minutes. Teams: pairs or groups of three. Prep: one worked hypothesis test or interval, written out with three deliberate mistakes.
Project the flawed write-up. Teams find and correct all three errors, one point for finding a mistake and a second point for explaining why it is wrong. Plant errors that mirror real ones: a condition borrowed from the wrong procedure, a standard error computed with the wrong denominator, and a p-value described as the probability the null hypothesis is true.
This is harder than producing a correct write-up from scratch, because spotting a wrong condition or a misread p-value requires knowing what the correct one looks like rather than following a memorized sequence. A full worked example, with all three errors and their fixes, is below. Trains: condition checking and p-value interpretation, the two most commonly missed lines. Fits: any inference unit, Unit 3 through Unit 5.
3. Predict, commit, reveal on the coverage simulator
Time: 8 minutes. Teams: whole class, no groups. Prep: none beyond opening the confidence interval coverage simulator.
Set the true proportion and the confidence level, then before drawing a single interval, ask the class to commit in writing to how many of the next 25 intervals will miss the true proportion. Draw the 25, count the misses on screen, and compare against each guess. Change the confidence level from 95 to 90 and run it again; ask what should happen to the miss count before revealing it.
Trains: the single most misquoted sentence on the exam, what 95 percent confidence actually means. Students remember a correction to a number they committed to and forget an explanation of a number they never guessed at. Fits: Unit 3 and Unit 4, wherever confidence intervals are taught or reviewed.
4. Kill the correlation
Time: 10 minutes. Teams: whole class or small groups taking turns at the keyboard. Prep: none beyond opening the regression influential point explorer.
Show the starting scatterplot and its correlation. Before anyone touches a point, the class predicts what will happen to r and to the sign of the slope if the flagged point is dragged straight down. One student drags it while the rest watch the least squares line and the residual plot move live, then the class checks its prediction against the number on screen.
A worked run of this exact game, with the before and after numbers, is below. Trains: why a single unusual point can matter more than every other point combined, which is the idea behind an influential point rather than a fact to memorize. Fits: Unit 5, regression.
5. Beat the test chooser
Time: 10 minutes. Teams: small groups, buzzer optional. Prep: a list of scenario stems, one variable or two, one sample or two, categorical or quantitative.
Read a scenario aloud. Groups call out z-test, t-test, chi-square, or regression before anyone opens a tool. Then walk the same scenario through the which statistical test interactive as a class and check the two or three questions it asks against the reasoning the groups just used out loud.
Trains: test selection, which is worth nothing on its own line of a rubric but silently caps every other point on an inference free response if it is wrong. Reading whether data is paired or independent, and whether the question is about a proportion or a mean, is the actual skill under the buzzer. Fits: Unit 3, Unit 4, and Unit 5 together, since the choice between them is the point.
6. Category speed round
Time: under a minute per item. Teams: whole class, hands or cards. Prep: a stack of quick scenarios.
Call a scenario or a pair of events. Students place it into a bucket in five seconds: independent or not independent, mutually exclusive or not, paired data or two independent samples, reject the null or fail to reject. Keep rounds short enough that instinct is being tested rather than a worked calculation.
Trains: the classification speed that the multiple choice section rewards, and it surfaces confusions fast, disjoint events getting called independent is the single most common one. Fits: Unit 2 for events, Unit 1 and Unit 3 for study design and paired versus independent samples.
7. Free response jigsaw
Time: 20 minutes. Teams: four groups, one part each. Prep: one released-style free response question, split at its natural parts.
Each group answers only its assigned part in full, then presents it to the class in order. Group two usually discovers that its part depends on group one's answer being right, and group four usually discovers that its conclusion has to reference numbers from parts it did not compute itself.
Trains: the fact that a four-part free response is one connected argument, not four unrelated questions with the same scenario at the top. That is the most useful realization a review lesson can produce. Fits: any unit, and it pairs naturally with the FRQ format guide for choosing which question to cut apart.
8. Two truths and a lie, statistics edition
Time: 10 minutes. Teams: pairs writing, whole class guessing. Prep: none, just paper.
Each pair writes three statements about one concept, two true and one false but plausible enough to fool the room. Swap with another pair, who has ninety seconds to identify the lie and say why. A weak lie is obviously wrong; a good one requires understanding the concept well enough to bend it convincingly.
Trains: writing the lie is where the real thinking happens. To make a convincing false statement about a sampling distribution, a student has to understand the true one well enough to distort a single detail. Fits: any vocabulary-heavy stretch, especially Unit 1 through Unit 3.
9. The sixty-second explanation
Time: one minute per turn. Teams: individuals, class watching. Prep: a stack of terms and, for each, three words to ban.
Draw a term at random and give the student sixty seconds to explain it to the class without using three banned words. For p-value, ban "significant," "probability," and "reject." For a confidence interval, ban "confident," "sure," and "percent."
Trains: genuine understanding over recited definitions. A student who only knows the textbook sentence cannot do this, and it becomes obvious to the room within about ten seconds, which is exactly the information a teacher needs two weeks out. Fits: any unit; the glossary is the fastest source of terms and their standard definitions to draw from.
10. The condition audit
Time: 12 minutes. Teams: pairs or groups of three. Prep: three or four short scenarios, each with a condition quietly violated or unverifiable.
Give each group a scenario, a sample size, and a claim about how the data were collected. Groups list every condition the chosen procedure requires, mark each as met, violated, or unverifiable from what is given, and state what to do about any violation, use a different procedure, note the limitation, or say the inference cannot proceed.
Plant scenarios where the sample size is under thirty with no shape described, where the 10 percent condition is not addressed for a sample that is not obviously small relative to the population, and where independence is assumed without random assignment or random sampling. Trains: the checklist itself, covered fully in conditions for inference, applied to cases where the answer is not automatically yes. Fits: Unit 3 and Unit 4, wherever a procedure has conditions attached to it.
Running a review day before the exam
A single review day works best run in three blocks rather than as one long practice test.
Morning, forty minutes: recall and classification. Category speed rounds, two truths and a lie, and sixty second explanations, cycling through vocabulary from every unit. These are fast, loud, and good for waking a room up, and they surface which terms need a full reteach before the day moves on.
Midday, fifty minutes: the model in motion. The State-Plan-Do-Conclude relay, the condition audit, kill the correlation, and beat the test chooser. These run slower and reward groups that talk through their reasoning rather than the fastest hand in the room.
Afternoon, about thirty-five minutes: the error hunt and the free response jigsaw. By this point most of the remaining gains are in writing rather than knowing, so close on formats that force a full written argument rather than a shout-out answer.
Two things to avoid on this specific day. Do not run a review day with no writing; the free response section is half the exam and it is graded on the sentence, not just the number. And do not let one confident student carry a group through every stage; the relay format fixes this on its own because the marker has to change hands.
Keep the formula sheet, the z-table, and the t-table open on a side screen the whole day, since the actual exam provides both, and a review day that never uses them teaches a habit the exam does not reward. The cram sheet is the fastest reference for pulling scenarios and numbers for any of the ten games above, and students who want an individually paced version of the same review afterward can run the study plan or the timed sets in practice.
Worked example: the error hunt on a one-sample t-test
A projected write-up claims to test whether a machine fills bottles to a mean of 100 milliliters. A sample of 25 bottles gives a mean of 97 mL and a standard deviation of 9 mL, tested at alpha equals 0.05. The projected write-up reads: State, H0: mu = 100, Ha: mu is not 100. Plan, one-sample t-test since n = 25 is under 30; Random given, 10 percent condition satisfied, Large Counts condition met since n is at least 10. Do, standard error equals s over n equals 9 over 25 equals 0.36; t equals (97 minus 100) over 0.36, equals negative 8.33, df = 24, p-value is approximately 0.0000004. Conclude, since the p-value is less than alpha, reject H0; there is a 99.9999 percent probability that the null hypothesis is false. Find and correct the three planted errors.
Error one, the condition. Large Counts is the condition for a proportion, not a mean. For a one-sample t-test with n under 30, the condition is that the population distribution is approximately normal, checked from a graph of the sample, such as a boxplot or dotplot showing no strong skew and no outliers.
Error two, the standard error. The formula divides the sample standard deviation by the square root of n, not by n itself. The correct value is 9 over the square root of 25, which is 9 over 5, or 1.8, not 9 over 25.
Recompute the test statistic with the corrected standard error. t equals (97 minus 100) over 1.8, which is negative 1.67, with 24 degrees of freedom.
Recompute the p-value from the corrected t. A t-statistic of negative 1.67 on 24 degrees of freedom gives a two-sided p-value of about 0.109, nowhere near the fabricated 0.0000004 that came from the inflated test statistic.
Error three, the conclusion. A p-value is the probability of a sample statistic this extreme or more, assuming the null hypothesis is true. It is never a probability that the null hypothesis itself is true or false. With the corrected p-value of about 0.109 greater than alpha equals 0.05, the correct decision is to fail to reject H0, the opposite of what the flawed write-up concluded.
The three errors are the condition name, Large Counts belongs to a proportion test and should be a normality check for a mean; the standard error, which should divide by the square root of n and equals 1.8, not 0.36; and the interpretation of the p-value, which is never a probability about the null hypothesis itself. Correcting all three gives t is about negative 1.67 on 24 degrees of freedom, a two-sided p-value of about 0.109, and a final decision of fail to reject H0, the reverse of the projected write-up's conclusion.
Worked example: kill the correlation on the influential point explorer
The starting scatterplot has nine points: (2, 4.2), (3.5, 5.4), (5, 6.1), (6, 7.6), (7.5, 8.1), (9, 9.8), (10.5, 10.4), (12, 11.9), and (18, 16.5), giving a correlation of about 0.998 and a positive least squares slope. The flagged point is the last one, at x equals 18, y equals 16.5, sitting right on the trend but far to the right of the other eight. The class is told this point will be dragged straight down, to y equals 1, with every other point held fixed. Predict what happens to r and to the sign of the slope, then check the prediction against the interactive.
Read the starting shape. All nine points, including the flagged point at the far right, sit close to one steady upward line, which is why the starting correlation is a near-perfect positive 0.998.
Identify why this one point matters more than the others. Even though it lies right on the trend, it sits far to the right of the other eight, at the largest x-value by a wide margin, which gives it outsized pull on where the least squares line tilts.
Predict the direction. Dragging it from the top of the pattern to the bottom removes the point anchoring the line's rise at the far right and replaces it with one pulling the line down at the same far edge, so both r and the slope should fall, likely past zero.
Check the prediction against the recomputed values. After the move, r drops to about negative 0.03 and the slope flips from about positive 0.77 to about negative 0.02, so the relationship goes from a near-perfect positive line to essentially flat and slightly negative from moving a single point.
State the lesson in one sentence. A correlation and a slope computed from nine points describe nine points, and one of them, sitting far out on the x-axis, was carrying almost the entire fit on its own even though it never looked unusual on the scatterplot itself.
Both r and the slope collapse. The correlation drops from about 0.998 to about negative 0.03, and the slope flips sign from about positive 0.77 to about negative 0.02, all from moving one of nine points. The lesson is that a correlation this size can rest almost entirely on a single far-out point, which is the reason a scatterplot has to be looked at before its summary number is trusted.
Frequently asked questions
What is the best AP Statistics review game?
The State-Plan-Do-Conclude relay. Four students share one whiteboard and one marker, each writing one part of a full inference write-up in order, and the format shows a teacher exactly which of the four parts a given team is weak at, which a finished answer alone does not reveal.
How do you review for AP Statistics without just doing more practice sets?
Use formats that force writing or reasoning instead of recognizing a correct option among four choices. The error hunt, where students find deliberate mistakes in a worked write-up, and the sixty second explanation with key words banned, both require understanding a procedure well enough to defend or bend it, which multiple choice practice does not test.
Are fast recall games useful for AP Statistics review?
Early in a unit, yes, category speed rounds and two truths and a lie build the vocabulary a later, slower game depends on. In the final two weeks, keep them short and pair them with formats that make students write a full sentence, since roughly half the exam is graded on the written argument rather than the classification alone.
What should the last review day before the AP Statistics exam look like?
Recall and classification first, while any weak vocabulary still needs shoring up, model-in-motion games such as the relay and the condition audit through the middle of the day once the procedures are back, and the error hunt and free response jigsaw last, when the remaining gains are in writing a full argument rather than knowing a fact.
How many students should be on a team for these games?
Three or four works for most of the games above. Anything larger lets one student carry the team through every stage, and the relay and jigsaw formats are built specifically to prevent that by forcing every member to hold the marker or own one part of the answer.