Describing and comparing distributions practice
By Jude Wallis · Published
This set has 8 problems on describing one distribution and comparing two. Each builds the same habit: name the shape, the center, the spread, and any unusual features, always in context, then compare with explicit comparative language. Solve each on paper before opening the steps.
AP Statistics: Unit 1 (topics 1.6 Descriptions for One Quantitative Variable Distributions, 1.9 Comparisons of the Distributions for One Quantitative Variable). These problems cover Unit 1 topics 1.6 (describing shape, center, spread, and unusual features for one quantitative variable) and 1.9 (comparing two such distributions) in the Fall 2026 AP Statistics course, where Unit 1 carries 20 to 30% of the multiple-choice section.
What these problems build
These 8 problems build the description workflow AP topics 1.6 and 1.9 score: shape, center, spread, and unusual features, each stated in terms of the variable being measured and its units. Shape covers the number of peaks and the direction of any skew: skewed right when the long tail runs toward the large values, skewed left when it runs toward the small ones, and bimodal when two clear peaks separated by a dip show up. Unusual features covers outliers, gaps, and clusters. Center is the median or the mean , and spread is the range, the interquartile range , or the sample standard deviation .
Problems 1, 2, 3, 5, and 6 describe a single distribution presented as a dotplot table, a raw list, a histogram, or a stemplot. Problems 4, 7, and 8 compare two distributions, working from a summary table, from raw lists, and from side-by-side boxplots. On a comparison question a correct calculation still earns nothing without comparative language, so those three are worth slowing down for. Nothing here is a hypothesis test: these are Unit 1 topics, so the work is describing what the data show, not deciding whether a difference is real. For the underlying methods, see skewed left vs skewed right, mean vs median, and how to find outliers with the 1.5 IQR rule. To check a five-number summary, open the five-number summary calculator, and to see how a graph reshapes as values move, try the descriptive statistics sandbox. More sets are on the practice page.
Comparative language, and the conventions used here
Two separate descriptions score nothing on a comparison question, even when both are correct. A response that says 'Group A has median 12 days and Group B has median 9 days' has described two distributions and compared neither. A response that says 'Group A's median germination time of 12 days is 3 days longer than Group B's median of 9 days' has compared them. The words doing the work link the two groups in a single sentence: higher than, lower than, more variable than, about the same as, more strongly skewed than. Compare all four elements, quote the numbers, and keep the units attached. One trap is worth planning for: the IQR and the range can point in opposite directions, so a group looks more consistent on one measure and less on the other. That is not a contradiction, it is usually an outlier, and the fix is to name which measure of spread you are using. See standard deviation vs IQR and dotplot vs histogram vs stemplot.
Which summaries to report follows from resistance: the mean and use every value, so one extreme observation drags both, while the median and the IQR depend on position rather than size. Report the mean and for a roughly symmetric distribution with no outliers, and the median and the IQR when it is clearly skewed or carries one. With no graph available, a mean well above the median suggests a right skew and a mean well below suggests a left skew, though the two sitting close together is only consistent with symmetry and does not prove it. Every quartile in these solutions uses the median-excluded (TI-84) convention: find the median first, then take as the median of the values strictly below it and as the median of the values strictly above it, leaving the median itself out of both halves when the count is odd. The sample standard deviation divides by .
Frequently asked questions
What order should I describe a distribution in?
Shape, center, spread, then unusual features is the usual order, and many students remember it as SOCS (shape, outliers, center, spread). Any order earns credit as long as all four appear and each one is tied to the variable and its units. The habit that loses points is naming a number without naming what it measures: is not a description, while 'the middle half of the finishing times spans 9 minutes' is.
When do I report the mean and standard deviation instead of the median and IQR?
Report the mean and when the distribution is roughly symmetric with no outliers, and the median and when it is clearly skewed or carries an outlier. The reason is resistance: the mean and the standard deviation are computed from every value, so one extreme observation drags both, as problem 2 shows when a single 2400-dollar listing pushes the mean above 8 of the 9 rents. If you are unsure, report the resistant pair and say why.
Why do my quartiles not match the ones my textbook gets?
Almost certainly a convention difference on an odd count. These solutions use the median-excluded (TI-84) rule: find the median, then take as the median of the values strictly below it and as the median of the values strictly above it, leaving the median value out of both halves. Some textbooks include the median in both halves instead, which shifts and . Pick one convention, use it consistently, and state it if a question leaves room for doubt.
Do I need a hypothesis test to say two distributions differ?
Not for this material. Topics 1.6 and 1.9 are descriptive, so the task is to report what the graphs and summaries show for the data in front of you. Say that one group's median is 2 days lower than the other's, not that the difference is statistically significant, and avoid claims about a wider population. Significance language belongs to the inference units, where you would have hypotheses, conditions, and a p-value behind it.
Problem 1
A teacher asks all 24 students in a homeroom how many pets live in their house and builds a dotplot. The counts are summarized here.
| Number of pets | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| Students | 5 | 8 | 6 | 4 | 0 | 0 | 0 | 1 |
Describe this distribution. Give the shape, a center, a measure of spread, and any unusual features, all in context.
Show the worked solution
Check the count first. students, which matches the homeroom size, so no responses are missing.
Shape. The tallest stack sits at 1 pet, and from there the stacks fall away steadily (8 students, then 6, then 4) before three empty values and a single dot far out at 7. One peak with a long thin tail toward the large values makes the distribution unimodal and skewed to the right.
Center. With the median is the average of the 12th and 13th values in order. The 5 students with 0 pets fill positions 1 through 5 and the 8 students with 1 pet fill positions 6 through 13, so both the 12th and the 13th values are 1, and the median is 1 pet. The mean is pets.
Spread. The counts run from 0 pets to 7 pets, a range of 7 pets. For the middle half, positions 1 through 12 are 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, so , and positions 13 through 24 are 1, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 7, so . Then pet.
Unusual features. No student reported 4, 5, or 6 pets, so there is a gap of three whole values before the largest observation. Testing it, , so the upper fence is pets, and the value 7 sits above it and is an outlier.
Put the four pieces together in context. The number of pets per household in this homeroom is unimodal and skewed to the right, centered at a median of 1 pet, with the middle 50% of students spanning just 1 pet () out of a full range of 7 pets. One student reported 7 pets, an outlier above the 3.5-pet fence and separated from everyone else by a gap at 4, 5, and 6 pets. Because the distribution is skewed and carries an outlier, the median of 1 pet describes a typical household better than the mean of 1.625 pets.
Unimodal and skewed right. Median 1 pet (mean 1.625 pets), pet, range 7 pets. Unusual features: a gap at 4, 5, and 6 pets, and one outlier at 7 pets, above the upper fence of 3.5.
Problem 2
A campus housing site lists the monthly rent, in dollars, for the 9 one-bedroom apartments within walking distance of a small college: 850, 900, 875, 950, 925, 880, 910, 890, 2400. A student blog reports that the average rent near campus is about 1064 dollars a month. Find the mean, the median, the standard deviation, and the IQR, then decide which pair of summaries describes these rents honestly and say why.
Show the worked solution
Sort the rents: 850, 875, 880, 890, 900, 910, 925, 950, 2400. There are listings.
Median. With 9 values the median is the 5th, so the median rent is 900 dollars.
Mean. The sum is , so dollars. That is the number the blog quoted.
Quartiles and IQR. The count is odd, so leave the median out of both halves. The lower half is 850, 875, 880, 890, giving , and the upper half is 910, 925, 950, 2400, giving . Then dollars.
Standard deviation. Working from the exact mean , the squared deviations sum to , so dollars.
Unusual features. , so the upper fence is dollars. The 2400-dollar listing sits far above that fence and is an outlier, while the other eight rents all fall within 100 dollars of one another.
Decide, and say why. One extreme listing makes the distribution strongly skewed to the right. It has pulled the mean up to 1064.44 dollars, which is higher than 8 of the 9 rents on the list, so the blog's average describes no apartment a student could actually rent. It has also inflated to 501.68 dollars, more than five times the whole 100-dollar spread of the other eight. Report the resistant pair instead: a typical one-bedroom near this campus rents for a median of 900 dollars, and the middle half of the listings span only 60 dollars ().
Mean 1064.44 dollars, median 900 dollars, dollars, dollars. The 2400-dollar listing is an outlier (upper fence 1027.5), so the distribution is strongly skewed right and the median and IQR are the honest summaries; the quoted mean exceeds 8 of the 9 actual rents.
Problem 3
A fitness app records the length, in minutes, of the 60 workouts one user logged last quarter. Each class includes its left endpoint and excludes its right endpoint.
| Length (min) | 0 to 10 | 10 to 20 | 20 to 30 | 30 to 40 | 40 to 50 | 50 to 60 | 60 to 70 |
|---|---|---|---|---|---|---|---|
| Workouts | 3 | 15 | 10 | 4 | 9 | 14 | 5 |
(a) Describe the shape. (b) Which class contains the median? (c) What percent of the workouts were shorter than 20 minutes? (d) Explain why a single measure of center is a poor summary of this distribution.
Show the worked solution
(a) Shape. Read the heights across the classes: 3, 15, 10, 4, 9, 14, 5. They climb to a peak of 15 at 10 to 20 minutes, fall to a low of 4 at 30 to 40 minutes, climb again to a second peak of 14 at 50 to 60 minutes, then fall away. Two clear peaks separated by a dip make the distribution bimodal. Side to side it is roughly balanced rather than skewed: the three classes below 30 minutes hold workouts and the three above 40 minutes hold .
(b) Median class. The total is workouts, so the median is the average of the 30th and 31st values in order. Running the cumulative counts gives 3 through the first class, 18 through the second, 28 through the third, and 32 through the fourth. Positions 29 through 32 therefore all sit in the 30 to 40 minute class, so both the 30th and the 31st values do, and the median lies in the 30 to 40 minute class.
(c) Percent under 20 minutes. The first two classes hold workouts, and , so 30% of the workouts were shorter than 20 minutes.
(d) Why one center misleads. The median lands in the 30 to 40 minute class, the shortest bar between the two peaks, holding only 4 of the 60 workouts. The center of this distribution falls in the valley between the two peaks, where almost nothing actually happened, so calling roughly 35 minutes a typical workout describes a length this user almost never trained for.
(d) Say what to do instead, in context. Two peaks usually mean two kinds of thing have been mixed into one graph, here plausibly short weekday sessions clustered around 10 to 20 minutes and long weekend sessions clustered around 50 to 60 minutes. The honest summary names both clusters and their centers rather than averaging across them.
(a) Bimodal, with peaks at 10 to 20 minutes and 50 to 60 minutes and a dip between, and roughly balanced side to side. (b) The 30 to 40 minute class, which holds the 30th and 31st values. (c) 30%. (d) The median falls in the dip between the peaks, a class holding only 4 of the 60 workouts, so no single center describes a typical workout; describe the two clusters separately.
Problem 4
Two farms each ship a random sample of 40 heirloom tomatoes to a grader, who weighs each one in grams and reports these summaries. No graph is available.
| Farm | Mean | SD | Min | Median | Max | ||
|---|---|---|---|---|---|---|---|
| A | 142 | 11 | 118 | 135 | 141 | 149 | 168 |
| B | 138 | 24 | 110 | 124 | 130 | 145 | 226 |
(a) What does each farm's mean-to-median relationship suggest about shape? (b) Which farm's weights are more variable? Support it with more than one measure. (c) Apply the 1.5 IQR rule to each farm. (d) Write a comparison in context.
Show the worked solution
(a) Shape from the mean against the median. Farm A: the mean of 142 g sits just 1 g above the median of 141 g, which is consistent with a roughly symmetric distribution. Farm B: the mean of 138 g sits 8 g above the median of 130 g, and a mean pulled well above the median is the signature of a right skew, almost certainly caused by a few unusually heavy tomatoes.
(b) Variability, three ways. Standard deviations: 11 g for Farm A against 24 g for Farm B. Interquartile ranges: g for A against g for B. Ranges: g for A against g for B. All three measures agree, so Farm B's weights are clearly the more variable.
(c) Outlier check, Farm A. g, so g. The fences are g and g. The minimum of 118 g and the maximum of 168 g both fall inside those fences, so the rule flags no outliers at Farm A.
(c) Outlier check, Farm B. g, so g. The fences are g and g. The minimum of 110 g is inside, but the maximum of 226 g is far above 176.5 g, so Farm B has at least one high outlier. A five-number summary cannot say how many tomatoes clear the fence, only that the heaviest one does.
(d) Comparison in context. The tomatoes sampled from Farm A are typically heavier than those sampled from Farm B, with a median weight of 141 g against 130 g, a difference of 11 g. The Farm A sample is also far more consistent, with an IQR of 14 g against Farm B's 21 g and a standard deviation of 11 g against 24 g. Farm A's weights look roughly symmetric with no outliers by the 1.5 IQR rule, while Farm B's are skewed to the right with at least one tomato above its 176.5 g fence, running as heavy as 226 g.
(a) Farm A's mean sits 1 g above its median (142 g vs 141 g), consistent with a roughly symmetric distribution; Farm B's sits 8 g above (138 g vs 130 g), which suggests a right skew. (b) Farm B, on SD (24 vs 11), IQR (21 vs 14), and range (116 vs 50). (c) Farm A has no outliers (fences 114 g and 170 g); Farm B has at least one high outlier, since 226 g exceeds the 176.5 g fence. (d) The Farm A sample is typically 11 g heavier at the median and much more consistent.
Problem 5
A biologist runs a trail camera and records the number of deer photographed each night for a week: 4, 6, 5, 7, 5, 6, 3. On the eighth night a herd passes and the camera records 26. (a) Find the mean, median, standard deviation, IQR, and range for the first 7 nights. (b) Recompute all five with the eighth night included. (c) Which summaries changed most, which barely moved, and what does that show about describing this camera's typical night?
Show the worked solution
(a) Sort the first seven counts: 3, 4, 5, 5, 6, 6, 7. The sum is 36, so deer. With 7 values the median is the 4th, so the median is 5 deer.
(a) Quartiles and range. Leaving the median out, the lower half is 3, 4, 5 and the upper half is 6, 6, 7, so and , giving deer. The range is deer.
(a) Standard deviation. The squared deviations from sum to , so deer.
(b) Add the eighth night. Sorted, the counts are 3, 4, 5, 5, 6, 6, 7, 26. The sum is 62, so deer, and with an even count the median is the average of the 4th and 5th values, deer.
(b) Quartiles, range, and standard deviation. The median falls between two values, so the lower half is 3, 4, 5, 5 and the upper half is 6, 6, 7, 26. Then and , so deer, exactly what it was before. The range jumps to deer. The squared deviations from 7.75 sum to , so deer.
(c) Line the two sets up. The mean climbed from 5.14 to 7.75 deer, a jump of 2.61 deer or about 51%. The standard deviation climbed from 1.35 to 7.48 deer, about 5.6 times as large. The range went from 4 to 23 deer. Against that, the median moved only from 5 to 5.5 deer and the IQR did not move at all.
(c) What it shows, in context. The mean and the standard deviation use every value, and the range uses the two extremes, so a single 26-deer night drags all three. The median and the IQR depend on position rather than size, so they absorbed the herd almost without flinching. Checking the rule on the eight nights, , so the upper fence is deer and the 26-deer night is an outlier. A description of what this camera sees on a typical night should therefore report the median of 5.5 deer with deer, and mention the 26-deer herd separately as an outlier rather than letting it set the average.
(a) Mean 5.14, median 5, , , range 4. (b) Mean 7.75, median 5.5, , , range 23. (c) The mean rose about 51%, became about 5.6 times as large, and the range went from 4 to 23, while the median moved 0.5 and the IQR did not change at all, showing the median and IQR are resistant and the mean, SD, and range are not.
Problem 6
Twenty runners finish a community 5K. Their times, in minutes, are shown in this stemplot, where a stem of 2 with a leaf of 4 means 24 minutes.
| Stem | Leaves |
|---|---|
| 1 | 8 9 |
| 2 | 0 1 2 3 4 5 5 6 7 8 9 |
| 3 | 0 1 2 4 6 7 |
| 4 | |
| 5 | 8 |
Describe the distribution of finishing times, and check the slowest time against the 1.5 IQR rule.
Show the worked solution
Read the times off the plot in order: 18, 19, 20, 21, 22, 23, 24, 25, 25, 26, 27, 28, 29, 30, 31, 32, 34, 36, 37, 58. That is runners, and the stemplot has already sorted them.
Shape. Eleven of the twenty leaves sit on the stem of 2, the count thins to six leaves on the stem of 3, the stem of 4 carries no leaves at all, and then one leaf appears on the stem of 5. The distribution is unimodal with its peak in the 20s and skewed to the right, with a gap across the entire 40s.
Center. With 20 values the median is the average of the 10th and 11th times, minutes. The mean is minutes. The mean sitting 1.75 minutes above the median is a second signal of the right skew.
Spread. The range is minutes. For the quartiles, the lower 10 times are 18, 19, 20, 21, 22, 23, 24, 25, 25, 26, so minutes, and the upper 10 are 27, 28, 29, 30, 31, 32, 34, 36, 37, 58, so minutes. Then minutes.
Outlier check. minutes, so the fences are minutes and minutes. The fastest time of 18 minutes is comfortably above the lower fence, but the slowest time of 58 minutes is well beyond the upper fence of 45 minutes, so it is an outlier.
Describe it in context. The finishing times for these 20 runners are unimodal and skewed to the right, centered at a median of 26.5 minutes, with the middle half of the field spanning 9 minutes () inside a full range of 40 minutes. One runner finished in 58 minutes, an outlier above the 45-minute fence and cut off from the rest of the field by an empty stem across the 40s. Because of that outlier, the median and the IQR describe this race better than the mean of 28.25 minutes and the standard deviation would.
Unimodal, skewed right, with a gap across the 40s. Median 26.5 minutes (mean 28.25 minutes), minutes, range 40 minutes. The 58-minute time is an outlier: it exceeds the upper fence of minutes.
Problem 7
A biology class tests whether a compost-enriched potting mix speeds up germination. Nine bean seeds go into standard mix and nine into enriched mix, and the class records the number of days until each seed sprouts.
- Standard mix: 6, 8, 7, 9, 7, 10, 8, 7, 12
- Enriched mix: 5, 6, 6, 7, 5, 8, 6, 7, 14
Compare the two distributions of germination time. Address shape, center, spread, and unusual features, and be careful with the two measures of spread.
Show the worked solution
Sort each group. Standard: 6, 7, 7, 7, 8, 8, 9, 10, 12. Enriched: 5, 5, 6, 6, 6, 7, 7, 8, 14. Each group has seeds.
Centers. With 9 values the median is the 5th. Standard has a median of 8 days and enriched has a median of 6 days. The means are days for standard and days for enriched.
Quartiles and spread, standard mix. Leaving the median out, the lower half is 6, 7, 7, 7 so , and the upper half is 8, 9, 10, 12 so . Then days and the range is days.
Quartiles and spread, enriched mix. The lower half is 5, 5, 6, 6 so , and the upper half is 7, 7, 8, 14 so . Then days and the range is days.
Unusual features. Standard: , so the upper fence is days, and the slowest seed at 12 days falls below it, so this group has no outliers. Enriched: , so the upper fence is days, and the seed that took 14 days is above it, so the enriched group has one high outlier.
Shape. In both groups the mean sits above the median (8.22 against 8, and 7.11 against 6) and the largest value stretches much farther above the median than the smallest falls below it, so both distributions are skewed to the right, the enriched group more strongly because of its outlier.
Resolve the conflicting spread measures before writing. The enriched group has the smaller IQR (2 days against 2.5 days) but the larger range (9 days against 6 days). Both are correct, and they are measuring different things: across the middle 50% of seeds the enriched mix was the more consistent, while its wider range comes entirely from the single 14-day seed. Name the measure whenever you make a spread claim.
Write the comparison in context. Seeds in the enriched mix germinated faster than seeds in the standard mix: the enriched median of 6 days is 2 days shorter than the standard median of 8 days. Both distributions are skewed to the right. Across the middle half of each group the enriched seeds were slightly more consistent, with an IQR of 2 days against the standard mix's 2.5 days, but the enriched group also holds one outlier at 14 days, above its 10.5-day fence, which stretches its range to 9 days against the standard mix's 6 days. The standard mix has no outliers by the 1.5 IQR rule.
Enriched median 6 days against standard median 8 days, so enriched germinated 2 days faster. Both are skewed right. Enriched days against standard 2.5 days, but enriched range 9 days against 6 days because of one outlier at 14 days (above its 10.5-day fence); the standard mix has no outliers (fence 13.25 days).
Problem 8
Two grocery stores in the same town report the hourly wages of their hourly employees. Store A has 32 such employees and Store B has 28. Side-by-side boxplots, drawn with whiskers running to the extremes rather than with the outlier rule applied, give these five-number summaries in dollars per hour.
| Store | Min | Median | Max | ||
|---|---|---|---|---|---|
| A | 15.00 | 16.50 | 18.00 | 20.00 | 24.00 |
| B | 14.00 | 17.00 | 21.00 | 23.50 | 26.00 |
(a) Which store pays more typically, and by how much? (b) Which store's wages are more variable? Support it with two measures. (c) About how many Store A employees earn between 16.50 and 20.00 dollars an hour? (d) An organizer claims that every Store B employee earns more than at least a quarter of Store A's employees. Evaluate the claim. (e) A second claim is that Store B's wages are bimodal. Can these boxplots settle it?
Show the worked solution
(a) Compare the medians, which is the measure of center a boxplot displays. Store B's median wage of 21.00 dollars an hour is higher than Store A's median of 18.00, so Store B typically pays dollars an hour more.
(b) Compare two measures of spread. Interquartile ranges: Store A has dollars and Store B has dollars. Ranges: Store A has dollars and Store B has dollars. Both measures point the same way, so Store B's wages are the more variable.
(c) The two values named, 16.50 and 20.00, are exactly Store A's and , and the box holds about the middle 50% of the data. So about 50% of Store A's 32 employees, roughly employees, earn between 16.50 and 20.00 dollars an hour.
(d) Evaluate the claim against the plot. Store B's minimum wage is 14.00 dollars an hour, which is below Store A's minimum of 15.00. So Store B's lowest-paid employee earns less than every single Store A employee, not more than a quarter of them, and the claim is false. One number settles it: Store B's left whisker starts to the left of Store A's.
(e) Bimodality. No, these boxplots cannot settle it. A boxplot is built from exactly five numbers, so it shows center, spread, and the rough direction of skew while hiding everything about the interior of the data. Peaks, gaps, and clusters are all invisible, and two very different distributions can produce identical boxplots. Deciding whether Store B's wages are bimodal, which would be plausible if the store pays one rate to cashiers and a higher rate to supervisors, needs a dotplot, a stemplot, or a histogram of the individual wages.
Note what else the display withholds. Neither mean can be recovered from a five-number summary, so nothing here supports a claim about average pay. And because the whiskers were drawn to the extremes, no point on either plot has been flagged as an outlier. Applying the rule by hand: Store A's fences are and dollars, and Store B's are and dollars. Every reported value sits inside its own store's fences, so neither store shows an outlier by the 1.5 IQR rule.
(a) Store B, by 3.00 dollars an hour at the median. (b) Store B: 6.50 against 3.50 and range 12.00 against 9.00. (c) About 16 employees. (d) False: Store B's minimum of 14.00 is below Store A's minimum of 15.00, so B's lowest earner is paid less than every Store A employee. (e) No: a boxplot shows only five numbers and hides peaks, gaps, and clusters, so a dotplot, stemplot, or histogram is needed.