AP Statistics Project Ideas for High School
By Jude Wallis · Published
Twelve high school statistics projects, each naming its deliverable, duration, the exact unit and topic it practices, and one rubric line, most also naming a free real data source, so the project reinforces the same skill the exam tests.
These twelve projects draw on skills from every unit of the course. Each one lists the exact topic code it practices, so a project can be assigned right alongside the unit that teaches the skill instead of only at the end of the course.
What makes a statistics project worth the class time
A statistics project earns its place when a student has to make a real decision under real uncertainty: which sampling method to use, whether a difference is large enough to call significant, whether a linear model actually fits. A poster with a pie chart on it does none of that.
The twelve projects below each name four things: the deliverable, how long it takes, the exact AP Statistics unit and topic code it practices, and a free source where the data is real rather than invented for the assignment. Most also name the one-line rubric criterion that catches the mistake students actually make on that project, not a generic completion check.
A note on data. "Real" does not have to mean downloaded. A tally a student takes standing in the cafeteria is real data with its own sampling problems to reason about, and several projects below use exactly that. Where an external source is named, it is free to access with no purchase or paid account required.
1. The sampling method audit
Duration: three to four days. Deliverable: a one-page audit of one real published survey, naming its sampling method and the specific bias that method risks.
Pew Research Center posts a full methodology statement with nearly every survey it releases, free to read, including how respondents were contacted and what the response rate was. Students pick one, identify whether the method was a simple random sample, a stratified sample, or something else, and then name the one bias that method is most exposed to rather than writing "there could be bias."
Rubric line: the sampling method is named correctly and the bias named is the one that method specifically risks, not a bias that applies to every survey.
Covers 1.11 Random Sampling and 1.12 Potential Problems with Sampling. Background in how to choose a sampling method and how to identify the type of bias.
2. The two-way table field count
Duration: one week, in three short observation sessions. Deliverable: a two-way table built from direct observation, such as tallying which lunch line students from each grade choose, with a bar graph and one sentence about whether the two variables look associated.
The data source is the school itself. Pick a location where two categorical variables can both be recorded on the spot, and record enough observations that no cell in the table stays empty.
Rubric line: every observed unit appears in exactly one cell, and the row and column totals sum to the same total observed.
Covers 2.1 Tabular and Graphical Representations for Two Categorical Variables. Sets up project 3 below.
3. The chi-square test of independence
Duration: two weeks. Deliverable: a full chi-square test of independence on a two-way table the student collects, with expected counts shown, the test statistic computed by hand at least once, and a written conclusion in the language of association rather than cause.
A workable version: survey classmates on grade level and typical mode of transportation to school, then test whether the two are independent. See the worked chi-square example below for exactly this setup with numbers filled in.
Rubric line: the conclusion sentence says "associated with" or "independent of," never "causes," and names the actual group surveyed rather than "students" in general.
Covers 3.14 and 3.15 Carrying Out a Chi-Square Test. Reference: chi-square tests explained and the chi-square calculator.
4. The stratified sample of the school
Duration: one week to design, one week to collect. Deliverable: a stratified random sample of the school population by grade level, used to estimate one real quantity such as average minutes of homework on a typical night, with the sample size in each stratum shown as proportional to that grade's share of the school.
Enrollment counts by grade are public for most public schools through the National Center for Education Statistics' free Common Core of Data lookup, or simply from the school's own front office.
Rubric line: the four strata do not overlap, and the number sampled from each is proportional to that stratum's actual share of total enrollment, not an equal split across grades.
Covers 1.11 Random Sampling. Background in simple random sample versus stratified sampling.
5. The confidence interval from a measured sample
Duration: one week. Deliverable: one confidence interval for a population mean, built from a quantity the class actually measures across a random sample of classmates: reaction time, resting pulse, or minutes of sleep the previous night. One interpretation sentence names the parameter and the confidence level.
Rubric line: the interpretation sentence states what the interval is confident about in context, such as "we are 95 percent confident the true mean reaction time of students at this school is between," not just the two numbers.
Covers 4.2 Constructing a Confidence Interval for a Population Mean. Steps in how to calculate a confidence interval; check the arithmetic with the confidence interval calculator.
6. The two-sample t-test on public sports statistics
Duration: one to two weeks. Deliverable: a two-sample t-test comparing a numeric statistic, such as points scored per game, between two seasons or two teams, with the standard error and degrees of freedom shown rather than only the final t-value and p-value. See the worked example below for a full version of this project.
Baseball-Reference, Basketball-Reference, and the rest of the Sports-Reference network publish complete free game logs and season statistics with no account required.
Rubric line: degrees of freedom and standard error appear as separate, labeled steps before the t-statistic, so a reader can check where an error entered.
Covers 4.9 and 4.10 Carrying Out a Test for the Difference Between Two Means. Background in pooled or not, two-sample t-test; check with the two-sample t-test calculator.
7. The correlation and regression project on economic data
Duration: one to two weeks. Deliverable: a scatterplot, correlation coefficient, least-squares regression equation, and a residual plot for two real quantitative variables tracked over time, with a written statement of whether the residual plot supports a linear model.
FRED, the Federal Reserve Bank of St. Louis's free public data site, publishes long, downloadable time series for hundreds of economic quantities with no account required.
Rubric line: the residual plot is shown and used to answer the linear-fit question; a high correlation coefficient alone does not earn full credit.
Covers 5.2 Correlation, 5.4 Residuals, and 5.5 Least-Squares Regression. Check with the regression calculator.
8. The completely randomized experiment
Duration: two to three weeks. Deliverable: a real experiment with random assignment on volunteers, such as two study methods measured against a quiz score, followed by a two-sample t-test on the results and a scope-of-inference statement about who the conclusion covers.
The data comes from the experiment itself. What makes this project hard to fake is that the random assignment mechanism has to be described precisely enough that another student could repeat it, not just the word "randomly."
Rubric line: the assignment method is described as an actual mechanism, such as shuffled cards or a random number generator applied to a numbered list, not the single word "randomly."
Covers 1.13 Experimental Design. Background in how to design an experiment and how to describe a completely randomized design; practice in experimental design practice.
9. The simulation project
Duration: three to four days. Deliverable: a simulation, run either with a random digit table by hand or with a spreadsheet's random function, estimating a real probability question, with the simulated estimate compared against a theoretical calculation when one exists.
A workable question: if a class of thirty has independent birthdays, what fraction of simulated classes will contain at least one shared birthday? Run one hundred trials and compare the simulated proportion to the theoretical probability.
Rubric line: the definitions of one trial and one success are written down before any trials are run, not reconstructed afterward to match the result.
Covers 2.3 Estimating Probabilities Using Simulation. See random digit table.
10. The sample size planning project
Duration: one week. Deliverable: a plan for the sample size needed to estimate a real campus proportion within a stated margin of error, computed before any data is collected, followed by an actual sample of that size and the resulting confidence interval.
Good proportions to estimate are ones a student can observe directly on campus: the share of cars in the lot with a parking permit, or the share of bins in a hallway with a recyclable item in the trash bin instead of the recycling bin.
Rubric line: the planned sample size is calculated and written down before data collection begins, and the actual interval width is compared with the target margin of error afterward.
Covers 1.12 Potential Problems with Sampling and 3.3 Constructing a Confidence Interval for a Population Proportion. Steps in how to find sample size for a margin of error; check with the sample size calculator.
11. The open data capstone
Duration: three to four weeks, run as an end-of-course project once every unit has been taught. Deliverable: a full report on one downloaded public dataset, covering distribution shape, summary statistics, one graph, and one inference procedure the student selects and justifies, with its conditions checked explicitly before running it.
OpenIntro publishes free data sets built specifically for introductory statistics classes, already documented with a description of each variable. Our World in Data publishes free, downloadable data on hundreds of real quantities tracked across countries and years.
Rubric line: the student states which inference procedure fits the data before running it and lists its conditions by name, rather than running a test first and checking conditions as an afterthought.
Covers material across every unit. Checklist in conditions for inference; practice in conditions checking practice.
12. The two-place comparison with a scope-of-inference statement
Duration: one to two weeks. Deliverable: a two-sample comparison, proportion or mean, of one characteristic between two real places, such as the unemployment rate in two counties or the median household income in two states, ending with a scope-of-inference sentence naming exactly who the data covers.
The U.S. Census Bureau's data site publishes free American Community Survey tables down to the county level, and the Bureau of Labor Statistics publishes free local unemployment and wage data.
Rubric line: the scope-of-inference sentence names the actual population the published data describes, not "everyone in the country."
Covers 3.9 through 3.13 on two-proportion inference. Reasoning in can you generalize these results; check with the proportion z-test calculator.
Grading twelve different projects without twelve different rubrics
The projects above cover five different units, but they can share one scoring structure, which is what actually keeps grading from becoming its own project. Score every project on the same four points:
| Part | Point |
|---|---|
| Names the correct procedure for the data type before running it | 1 |
| States and checks the conditions that procedure requires | 1 |
| Shows the calculation with each step labeled, not just a final number | 1 |
| Writes the conclusion in context, with the correct scope | 1 |
That structure repeats across a chi-square test, a t-test, and a regression, so a grader is checking the same four things every time rather than relearning a new rubric per project.
A few habits protect the time this saves. Require one named, real data source on every project; it is the cheapest check against a fabricated data set and it is also what makes the work statistics rather than arithmetic on invented numbers. Set the significance level once, for the whole class, before any project is turned in, so nobody can pick the threshold that makes their own result look better. Do not grade whether the conclusion "came out right"; a correctly run test that fails to reject the null hypothesis is not a worse project than one that rejects it, and grading toward a preferred outcome teaches students to shop for a result instead of running the procedure honestly. And skip the poster. A printed board tests layout, not the four points in the table above.
Chi-square test for independence: grade level and transportation mode
A student surveys 120 classmates on grade level (freshman or senior) and usual mode of transportation to school (bus, drive, or walk), to test whether grade level and transportation mode are independent. Observed counts: freshmen are 30 bus, 5 drive, 25 walk; seniors are 10 bus, 40 drive, 10 walk. Run the chi-square test of independence at the 5 percent significance level.
Find the row, column, and grand totals. Freshmen total 60, seniors total 60. Bus totals 40, drive totals 45, walk totals 35. The grand total is .
Compute each expected count as (row total times column total) divided by the grand total. Freshman bus: . Freshman drive: . Freshman walk: . By symmetry the senior row has the same three expected counts: 20, 22.5, and 17.5.
Check the three conditions before computing the statistic. Random: the 120 respondents are the student's own classmates, a convenience sample rather than a random sample of the whole school, so this condition is not met and the conclusion will need to stay scoped to the group actually surveyed. Independence: the sample is well under 10 percent of the school's total enrollment, and one respondent's answer does not affect another's, so individual observations can be treated as independent. Large Counts: every expected count is at least 5, the smallest being 17.5, so this condition is met. With Independence and Large Counts satisfied, the calculation can proceed, carrying that Random limitation into the conclusion below.
Compute for each of the six cells: freshman bus ; freshman drive ; freshman walk ; senior bus ; senior drive ; senior walk .
Sum the six values: . Degrees of freedom equal (rows minus 1) times (columns minus 1), or .
Compare to the critical value. At 2 degrees of freedom and a 5 percent significance level, the critical chi-square value is 5.991. The computed statistic, about 43.65, is far larger, so the p-value is well below 0.05.
State the conclusion in context and in scope. Reject the claim of independence: grade level and transportation mode appear associated among the students surveyed. Because the 120 respondents were the student's own classmates rather than a random sample of the whole school, the conclusion should be stated as applying to that surveyed group, not generalized to every student at the school.
Chi-square statistic approximately 43.65 on 2 degrees of freedom, well above the 5.991 critical value at the 5 percent level. There is convincing evidence of an association between grade level and transportation mode among the students surveyed. Because the sample was not randomly drawn from the full school, the conclusion should not be generalized beyond the group actually surveyed.
Two-sample t-test on season point totals from a public game log
A student pulls points scored per game for a random sample of 15 games from one season and 15 games from a different season of the same team, both from a free public game log. Season A: mean 102.4 points, standard deviation 8.1. Season B: mean 96.7 points, standard deviation 7.4. Test whether the mean points per game differ between the two seasons at the 5 percent significance level.
State the hypotheses. The null hypothesis is that the two seasons have the same true mean points per game; the alternative is that they differ.
Check conditions. Each sample of 15 games was randomly selected from its season's full game log, the two seasons are independent of each other, and with sample sizes of 15 the underlying distribution of points per game should be reasonably symmetric for the t-procedures to be appropriate; a quick look at each season's dotplot would confirm no strong skew or outliers.
Compute the standard error of the difference in means: .
Compute the t-statistic: .
Find the p-value using the conservative degrees of freedom, the smaller sample size minus 1, so . At , a t-statistic of about 2.01 falls between the values that mark two-sided p-values of 0.10 and 0.05, giving a two-sided p-value of roughly 0.06.
Compare the p-value to the significance level. Since roughly 0.06 is greater than 0.05, the test fails to reject the null hypothesis at the 5 percent level.
t is approximately 2.01 with 14 degrees of freedom and a two-sided p-value of roughly 0.06. At the 5 percent significance level, the sample does not provide convincing evidence that the mean points per game differ between the two seasons, even though the 5.7 point gap looks meaningful without a test attached to it.
Frequently asked questions
What is a good free source of real data for a high school statistics project?
It depends on the project. For public opinion data with full methodology attached, Pew Research Center. For economic time series, FRED from the Federal Reserve Bank of St. Louis. For sports statistics, the Sports-Reference network. For data built specifically for an intro statistics class, OpenIntro's free data sets. All of these are free to access with no purchase required, and several projects above work just as well on data a student collects directly at school.
Do statistics projects need a formal inference procedure, or is a graph enough?
A graph and summary statistics are a fine first project early in the course, matching Unit 1 skills. Once a class has reached confidence intervals or significance tests, a stronger project asks the student to choose and run the correct one, since deciding which procedure fits a real question is a large share of what the exam actually tests.
How do you grade twelve different projects without writing twelve different rubrics?
Use one four-point structure for all of them: the correct procedure named before running it, its conditions checked, the calculation shown step by step, and the conclusion written in context with the correct scope. That structure fits a chi-square test, a two-sample t-test, and a regression equally well, so grading time does not grow with the number of distinct projects.
Should a statistics project be graded on whether the result is significant?
No. A correctly run test that fails to reject the null hypothesis is not a weaker project than one that rejects it. Grading toward a preferred outcome teaches students to keep collecting data or adjusting the test until they get the answer they want, which is the opposite of the honesty the procedure is supposed to enforce.
Is the stock market a good source of data for a statistics project?
It works better for regression and correlation than for causal claims, since nobody assigned which stocks moved. A cleaner project pulls two real quantitative variables from a source such as FRED or a public sports log and lets the class practice scatterplots, correlation, and residual checking without implying that one variable caused the other.