Chi-square two-way table practice problems
By Jude Wallis · Updated
These 8 problems run chi-square tests on two-way tables end to end: expected counts for every cell, the expected-counts condition, the statistic term by term, degrees of freedom, and a conclusion in context. Two turn on which test the design calls for, and one has a condition that fails.
AP Statistics: Unit 3 (topics 3.14 Setting Up a Chi-Square Test for Homogeneity or Independence, 3.15 Carrying Out a Chi-Square Test for Homogeneity or Independence). These problems cover Unit 3 topics 3.14 and 3.15 of the Fall 2026 AP Statistics course, and Unit 3 carries 15 to 25% of the multiple-choice section. The chi-square goodness-of-fit test was removed in that redesign, so it appears here only as a distractor to recognize and rule out.
What these problems build
These 8 problems run a chi-square test on a two-way table of counts from start to finish: the expected count in every cell, the expected-counts condition, the statistic built one term at a time, the degrees of freedom, and a conclusion stated in context.
The expected count for a cell is , the count the null model predicts. The statistic is
where (read 'ky-square') sums over every cell, is an observed count, and is that cell's expected count. Degrees of freedom come from the shape of the table, , and not from the sample size.
Two of the problems turn on which test the design calls for. A test of homogeneity compares one categorical variable across two or more separate samples or treatment groups, while a test of independence asks whether two variables are associated inside a single sample. Read how the data were collected, not what the question is about. One problem has an expected count below 5 so you can practice reporting a failed condition instead of a p-value.
A third chi-square test, goodness-of-fit, compares one sample against a claimed set of proportions. It was removed from the AP Statistics exam in the Fall 2026 redesign, while homogeneity and independence both stayed, so it shows up here only as something to recognize and rule out. It is still standard content in a college introductory course.
For the setup and the conditions, read chi-square tests explained, and look up cutoffs on the chi-square table. Check your arithmetic against the chi-square calculator.
If you are unsure which procedure a scenario needs, work through which statistical test to use or the which test interactive. The AP topics behind this set are 3.14, setting up a chi-square test and 3.15, carrying one out. More sets are on the practice page.
Problem 1
A community garden association takes a random sample of 240 plot holders from its county-wide membership and records the size of the plot each one rents and whether that person uses the shared compost bins.
| Plot size | Uses compost | Does not | Total |
|---|---|---|---|
| Small | 68 | 22 | 90 |
| Large | 92 | 58 | 150 |
| Total | 160 | 80 | 240 |
Find the expected count for all four cells under the null hypothesis of no association, then check the expected-counts condition.
Show the worked solution
Every expected count is . The row totals are 90 (small) and 150 (large), the column totals are 160 (uses compost) and 80 (does not), and the grand total is 240.
Small and uses compost: .
Small and does not: .
Large and uses compost: .
Large and does not: .
Check the expected counts against the margins. They add to across the small row and across the large row, matching the observed row totals, which is a fast way to catch an arithmetic slip.
Check the condition. The four expected counts are 60, 30, 100, and 50, and every one is greater than 5, so the expected-counts condition is met.
Expected counts: small and compost 60, small and no compost 30, large and compost 100, large and no compost 50. All four exceed 5, so the expected-counts condition is met.
Problem 2
For each study, name the chi-square test that fits and give the degrees of freedom it would use.
- Study A. A parks department draws independent random samples of 120 swimmers at each of four public pools and records which session each swimmer prefers: morning, afternoon, or evening.
- Study B. A librarian takes one random sample of 400 card holders and records both the genre each person borrows most (fiction, nonfiction, graphic novels, or audio) and whether that person uses the library app.
- Study C. A theater takes one random sample of 300 ticket sales and records only the showtime purchased, then asks whether the four showtimes sell in the 40%, 30%, 20%, 10% split the owner claims.
Show the worked solution
Read the design, not the topic. Ask two questions about each study: how many populations were sampled, and how many variables were measured on each subject.
Study A. Four separate samples, one per pool, each measured on a single categorical variable (preferred session). Comparing the distribution of one variable across several populations is a chi-square test for homogeneity. The table is 4 pools by 3 sessions, so .
Study B. One sample from one population, measured on two categorical variables (genre and app use). Asking whether the row and column variables are associated inside that single population is a chi-square test for independence. The table is 4 genres by 2 app categories, so .
Study C. One sample again, but only one variable, compared against a claimed set of proportions rather than against another group. That is a chi-square goodness-of-fit test. There is no two-way table, so the degrees of freedom are the number of categories minus 1, or .
A note on Study C for AP students. The goodness-of-fit test was removed from the AP Statistics exam in the Fall 2026 redesign, while the tests for homogeneity and for independence both remain. It is still standard in a college introductory course, so recognizing it is worth your time, but no AP question will ask you to carry it out.
A: homogeneity (four samples, one variable), . B: independence (one sample, two variables), . C: goodness-of-fit (one sample, one variable against a claimed distribution), , and that test is no longer on the AP exam.
Problem 3
A bakery wants to know whether crust quality depends on when a sourdough loaf is baked. It scores a random sample of 120 loaves from the morning bake and a separate random sample of 120 loaves from the evening bake as either pass or fail on a crust check.
| Bake | Pass | Fail | Total |
|---|---|---|---|
| Morning | 84 | 36 | 120 |
| Evening | 66 | 54 | 120 |
| Total | 150 | 90 | 240 |
Compute the chi-square statistic term by term, give the degrees of freedom, and decide at using the critical value 3.841.
Show the worked solution
Name the test. Two separate samples compared on one categorical variable is a chi-square test for homogeneity. : the pass rate on the crust check is the same for the two bakes. : the pass rate is not the same for the two bakes.
Expected counts. Both row totals are 120 and the grand total is 240, so each expected count is half its column total: in the pass column and in the fail column, for each bake. All four exceed 5.
Morning terms. Pass: . Fail: .
Evening terms. Pass: . Fail: .
Add the four terms. .
Degrees of freedom. The table is 2 by 2, so .
Decision. Since , reject at ; the p-value is about 0.016. There is convincing evidence that the pass rate on the crust check differs between the morning and evening bakes.
Expected counts 75 pass and 45 fail for each bake; with . Since (p about 0.016), reject : the pass rates differ.
Problem 4
A veterinary clinic takes a random sample of 80 exotic-pet owners from its client list and records the animal group each one owns and whether that owner carries pet insurance.
| Animal group | Insured | Not insured | Total |
|---|---|---|---|
| Reptile | 8 | 32 | 40 |
| Bird | 5 | 7 | 12 |
| Small mammal | 7 | 21 | 28 |
| Total | 20 | 60 | 80 |
(a) Find the expected counts and check the expected-counts condition. (b) Say what you may and may not report. (c) The clinic decides to combine the bird and small-mammal rows into one non-reptile row. Redo the check and carry out the test at .
Show the worked solution
(a) Expected counts, . Reptile: insured and not insured. Bird: insured and not insured. Small mammal: insured and not insured.
(a) Check the condition. Every expected count must be greater than 5, and the bird-and-insured cell has an expected count of 3, so the condition fails. Notice that the observed count in that cell is 5; the condition is checked on expected counts only, never on observed ones.
(b) What you may report. You can still describe the table with percentages, and the arithmetic for the statistic still runs, but the chi-square distribution is a poor approximation when an expected count is that small, so any p-value or critical-value decision read off would be untrustworthy. Report the failed condition rather than a p-value.
(c) Combine and recheck. The non-reptile row is insured and not insured, for a row total of 40. The table is now 2 by 2 with both row totals equal to 40, so every cell in the insured column expects and every cell in the other column expects . All four exceed 5, so the condition is met.
(c) Chi-square term by term. Reptile insured: . Reptile not insured: . Non-reptile insured: . Non-reptile not insured: .
(c) Statistic, degrees of freedom, decision. with . Since , and the p-value is about 0.302, fail to reject : there is not convincing evidence that insurance status is associated with whether the pet is a reptile.
A caution about combining categories. Merging is honest only when the merged group means something on its own and the decision is made for subject-matter reasons, ideally before anyone sees the counts. Merging categories because that merge produced the answer you wanted is data dredging, and collecting more data is the cleaner fix.
(a) Expected counts 10, 30, 3, 9, 7, 21; the bird-and-insured expected count of 3 is below 5, so the condition fails. (b) Report the failed condition, not a p-value. (c) Combined 2 by 2: expected 10 in each insured cell and 30 in each not-insured cell, , , and , so fail to reject .
Problem 5
A coding bootcamp randomly assigns 360 accepted applicants to three delivery formats, 120 to each, and records at the end whether each student completed the program.
| Format | Completed | Withdrew | Total |
|---|---|---|---|
| In person | 102 | 18 | 120 |
| Hybrid | 90 | 30 | 120 |
| Online | 78 | 42 | 120 |
| Total | 270 | 90 | 360 |
(a) Name the test and state the hypotheses. (b) Check the conditions. (c) Compute the statistic and the degrees of freedom. (d) Conclude at .
Show the worked solution
(a) Test and hypotheses. Three groups created by random assignment are compared on one categorical response, so this is a chi-square test for homogeneity across treatments. : the distribution of completion status is the same for all three formats. : the distribution of completion status is not the same for all three formats.
(b) Conditions. Randomization holds because students were randomly assigned to the three formats. The 10% condition does not apply here, since this is a randomized experiment rather than sampling without replacement from a population. The expected counts are checked in the next step.
(c) Expected counts, . Every row total is 120, so each format expects completions and withdrawals. All six exceed 5, so the expected-counts condition is met.
(c) Completed terms. In person: . Hybrid: . Online: .
(c) Withdrew terms. In person: . Hybrid: . Online: .
(c) Statistic and degrees of freedom. , and .
(d) Conclusion. The critical value for at is 9.210, and ; the p-value is about 0.0017. Reject : there is convincing evidence that completion rates differ across the three delivery formats. Because formats were randomly assigned, the difference can be attributed to the format itself.
(a) Homogeneity across three treatments. (b) Random assignment holds, the 10% condition is not needed, and expected counts of 90 and 30 all exceed 5. (c) , . (d) Since (p about 0.0017), reject ; completion rates differ by format, and random assignment supports a causal reading.
Problem 6
A county energy office takes one random sample of 300 owner-occupied homes and records each home's primary heat source and the era in which the home was built.
| Heat source | Before 1980 | 1980 to 2009 | 2010 or later | Total |
|---|---|---|---|---|
| Gas | 66 | 30 | 24 | 120 |
| Electric | 42 | 36 | 22 | 100 |
| Wood | 42 | 24 | 14 | 80 |
| Total | 150 | 90 | 60 | 300 |
Carry out a chi-square test for independence at , then interpret the result carefully.
Show the worked solution
Test and hypotheses. One random sample measured on two categorical variables calls for a chi-square test for independence. : primary heat source and build era are independent among owner-occupied homes in this county. : primary heat source and build era are associated.
Expected counts, . Gas row: , , and . Electric row: , , and . Wood row: , , and . All nine exceed 5.
Gas terms. , , and .
Electric terms. , , and .
Wood terms. , , and .
Statistic and degrees of freedom. , and .
Decision. The critical value for at is 9.488, and ; the p-value is about 0.327. Fail to reject .
Careful interpretation. There is not convincing evidence that heat source and build era are associated among owner-occupied homes in this county. That is not the same as showing the two variables are independent, because 300 homes spread across nine cells may simply be too few to detect a modest association.
with . Since (p about 0.327), fail to reject : no convincing evidence of an association, which is not the same as evidence that the two variables are independent.
Problem 7
A county extension office takes one random sample of 300 shoppers at a farmers market and records how far each shopper traveled and the category that shopper spent the most on.
| Distance | Produce | Prepared food | Crafts | Total |
|---|---|---|---|---|
| Under 2 miles | 42 | 12 | 6 | 60 |
| 2 to 10 miles | 54 | 42 | 24 | 120 |
| Over 10 miles | 54 | 36 | 30 | 120 |
| Total | 150 | 90 | 60 | 300 |
(a) Compute and the degrees of freedom. (b) Decide at and state a conclusion in context. (c) Identify the cell that contributes most and read what it says. (d) Interpret the p-value in one sentence.
Show the worked solution
(a) Expected counts, . Under 2 miles: , , and . Each of the two 120-shopper rows: , , and . All nine exceed 5.
(a) Under 2 miles terms. , , and .
(a) 2 to 10 miles terms. , , and .
(a) Over 10 miles terms. , , and .
(a) Statistic and degrees of freedom. , and .
(b) Decision and conclusion. The critical value for at is 9.488, and ; the p-value is about 0.0091. Reject : there is convincing evidence that distance traveled and main spending category are associated among shoppers at this market. The statistic also clears the cutoff of 13.277, so the same decision holds at the stricter level.
(c) Largest contribution. The nearest shoppers buying produce contribute 4.8, more than any other cell. Those shoppers bought produce 42 times when independence predicts 30, so the association is driven most by people who live within walking distance treating the market as a grocery run.
(d) Interpreting the p-value. If distance traveled and spending category were truly independent among this market's shoppers, only about 0.9% of random samples of 300 shoppers would produce a chi-square statistic of 13.5 or larger.
(a) , . (b) (p about 0.0091), so reject : distance and spending category are associated at this market. (c) The under-2-miles produce cell contributes 4.8 (42 observed against 30 expected). (d) Under independence, about 0.9% of samples of 300 would give a statistic this large or larger.
Problem 8
A gym chain wants to know whether one-year renewal rates are the same at its three locations. It takes independent random samples of 240 members from each location and reports only the renewal percentages: Riverside 85%, Downtown 75%, and Northgate 65%.
(a) Name the test and explain what in the design points to it. (b) Rebuild the two-way table of counts. (c) Check the expected-counts condition. (d) Compute and the degrees of freedom, then conclude at . (e) A manager reads the result and says it proves the Riverside location causes members to renew. Explain what is wrong with that.
Show the worked solution
(a) Test and why. Three separate random samples, one drawn from each location, are compared on one categorical variable (renewed or not). Comparing one variable's distribution across several populations is a chi-square test for homogeneity; a test for independence would require a single sample measured on two variables. : the distribution of renewal status is the same at all three locations. : it is not the same at all three.
(b) Rebuild the counts. Renewals: Riverside , Downtown , Northgate . Non-renewals are the rest of each 240: , , and .
(b) Totals. Renewed: . Did not renew: . The grand total is , which matches .
(c) Expected counts, . Each row total is 240, so every location expects renewals and non-renewals. All six exceed 5, and the samples are random and independent, so the conditions are met.
(d) Renewal terms. Riverside: . Downtown: . Northgate: .
(d) Non-renewal terms. Riverside: . Downtown: . Northgate: .
(d) Statistic, degrees of freedom, conclusion. with . Since and the p-value is well under 0.001, reject : there is convincing evidence that renewal rates differ across the three locations.
(e) Why the causal claim fails. Members were not randomly assigned to locations; they chose them, so location is tangled up with commute, income, and everything else that varies by neighborhood. A chi-square test on observational data can show that the distributions differ, but it cannot say the location caused the difference.
(e) One further limit. The test is a single overall comparison, so rejecting the null says the three distributions are not all the same, not which pair differs. The cell contributions of 9.6 at both Riverside and Northgate point to where the gap sits, but they describe the table rather than testing any pair formally.
(a) Homogeneity: three separate samples, one categorical response. (b) Renewed 204, 180, 156; not renewed 36, 60, 84; grand total 720. (c) Expected 180 and 60 at every location, all above 5. (d) , , p well under 0.001, so reject . (e) Members chose their location, so with no random assignment the test shows a difference, not a cause.