Dotplot, stemplot, and histogram practice
By Jude Wallis · Published
This set has 8 problems on the three graphs for one quantitative variable, dotplots, stemplots, and histograms, plus the boxplot as a picture of summary statistics. Four turn on choices that change what a graph shows, bin width above all. Solve each on paper before opening the steps.
AP Statistics: Unit 1 (topics 1.5 Graphical Representations for One Quantitative Variable, 1.8 Graphical Representations of Summary Statistics for One Quantitative Variable). These problems cover Unit 1 topic 1.5 (dotplots, stem-and-leaf plots, and histograms for one quantitative variable) and topic 1.8 (the boxplot, the graph built from the five-number summary) in the Fall 2026 AP Statistics course. Unit 1 carries 20 to 30% of the multiple-choice section.
What these problems build
These 8 problems cover AP topic 1.5, the graphical representations for one quantitative variable, and topic 1.8, the graph built from summary statistics. Problems 1, 4, and 5 use dotplots and stemplots, which keep every original value. Problems 2, 3, 6, and 8 use histograms, which keep only how many values fell in each bin. Problem 7 asks what a boxplot loses by keeping five numbers and nothing else.
The skill worth the most here is knowing that a graph is a choice, not a fact. Problem 2 takes one set of 30 waiting times and bins it twice: at width 10 it is unimodal, at width 5 it is bimodal, and neither picture is a mistake. Problem 3 flips the rule for values landing exactly on a bin edge and moves 6 of 20 test scores into a different bar, changing which bar is tallest. Problem 4 splits stems and lifts a gap into the outline of the rows, where the unsplit plot had left it visible only in the leaves. In every case the data never moved.
For the underlying methods, see how to read a stemplot, histogram vs bar graph, how to make a boxplot, and the bin width entry. The three displays are lined up side by side in dotplot vs histogram vs stemplot. To watch a histogram reshape as values move, open the descriptive statistics sandbox. More sets are on the practice page.
Conventions used here, and what these problems are not
Bin boundaries. Unless a problem says otherwise, every bin includes its left endpoint and excludes its right, so a value of 80 goes in the 80-to-90 bin rather than the 70-to-80 bin. Nothing drawn on a histogram records which rule was used, which is the whole of problem 3, so a display has to state its rule in words.
Stemplot keys. Every stemplot needs a key, because a row reading 6 with a leaf of 3 could mean 63, 6.3, or 630. Splitting stems, writing each stem twice with low leaves 0 to 4 and high leaves 5 to 9, halves the width of a row from 10 units to 5 and is the same move as narrowing a histogram's bins. Include stems that end up with no leaves, since the empty row is what makes a gap visible.
Quartiles. All quartiles use the median-excluded (TI-84) convention: find the median, then take as the median of the values strictly below it and as the median of the values strictly above it, leaving the median itself out of both halves when the count is odd. The sample standard deviation divides by .
Scope. Writing a full description in context, naming shape, center, spread, and unusual features, is drilled in describing distributions practice, and the quartile and outlier arithmetic is drilled in boxplot and five-number summary practice. The question in this set is narrower: what does this particular display show, what does it throw away, and what would a different display of the same numbers have shown instead.
Frequently asked questions
How do I choose a bin width for a histogram?
Start from the range divided by the number of bins you want, then round to a readable number. In problem 2 the range is 45 minutes, so eight bins suggests , rounded to 5. Then try a second width before you describe the shape. Too few bins hides structure, as problem 2 shows when one wide bar merges two clusters, and too many turns single observations into spikes until every bar is 0 or 1 tall and the shape is gone.
Which bin does a value on a boundary go in?
Whichever the display says. The usual convention, and the one used throughout this set, puts a value in the bin whose lower bound it meets or exceeds and whose upper bound it falls below, so 80 goes in the 80-to-90 bin. The other convention exists and problem 3 shows it moving 6 of 20 scores and changing the tallest bar. Nothing in the drawn picture records the choice, so state it. If you are picking the edges yourself, put them where no data value can land.
When is a stemplot better than a histogram?
When the data set is small enough to write out and the exact values matter. A stemplot keeps every digit, so you can read the median, the quartiles, the extremes, and any individual value straight off it, which problem 5 uses to get an exact median of 80.5 degrees where the histogram of the same data gives only 'between 80 and 90'. Roughly 15 to 50 values is the readable range. Past that the leaves run off the page and a histogram is the better summary.
Can two data sets have the same boxplot but different shapes?
Yes, and problem 7 builds a pair. A boxplot draws five numbers, so any two data sets sharing a five-number summary produce the same picture no matter what happens between those five points. The two trays there also share a mean of 20 cm, yet one is close to uniform and the other is bimodal with a near-empty middle, and their standard deviations differ (6.63 cm against 7.22 cm). Pair a boxplot with a dotplot or histogram whenever gaps, clusters, or extra peaks would change your conclusion.
Why can I read a median off a dotplot but not a histogram?
A dotplot puts one dot per observation at its own value, so the whole data set can be read back off it and counting to the middle gives the median exactly. A histogram replaces the values in each interval with a single count, so those values are gone: it can tell you which bin the median falls in, as in problem 8, but not where inside that bin. The same is true of the mean, the quartiles, and any question about a threshold that falls inside a bin.
Problem 1
Twenty-five students each add songs to a shared class playlist. A dotplot of the number of songs per student has these stack heights.
| Songs added | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Students (dots) | 3 | 8 | 6 | 4 | 2 | 1 | 0 | 1 |
(a) How many students are in the data set, and what is the largest number of songs one student added? (b) Find the mode, the median, and the mean. (c) A classmate says the tallest stack marks the center of the distribution. Correct that. (d) Redraw the same data as a histogram with bins of width 3 starting at 0, so the bins are 0 to 3, 3 to 6, and 6 to 9, each including its left endpoint and excluding its right. Give the three bar heights and say what the histogram hides.
Show the worked solution
(a) Every dot is one student, so add the stack heights: students. The rightmost dot sits above 8, so the largest count is 8 songs. Note also that the position at 7 songs is empty, a gap of one whole value before that last student.
(b) Mode. The tallest stack is the one at 2 songs, with 8 students, so the mode is 2 songs.
(b) Median. With the median is the 13th value in order. The cumulative counts are 3 students through 1 song, through 2 songs, and through 3 songs, so positions 12 through 17 all sit at 3 songs. The 13th value is therefore 3, and the median is 3 songs.
(b) Mean. Weight each value by its stack height: songs.
(c) Correct the claim. A stack height counts how often a value occurred, so the tallest stack is the mode, not the center. Center is a position in the ordered data, and here the three summaries separate: the mode is 2 songs, the median is 3 songs, and the mean is 3.08 songs. The thin right tail out to 8 songs pulls the mean above the median, and the mode sits below both.
(d) Bin the same 25 values. The bin from 0 to 3 holds the values 0, 1, and 2, so its height is . The bin from 3 to 6 holds 3, 4, and 5, so its height is . The bin from 6 to 9 holds 6, 7, and 8, so its height is . Check the total: .
(d) What the histogram hides. Its tallest bar is now the middle one, covering 3 to 6 songs, while the single most common value, 2 songs, is buried inside the shorter first bar. The gap at 7 songs disappears too, because 7 shares a bin with 6 and 8. Binning traded the individual values for three counts, and at this width it also moved the visible peak.
(a) 25 students, and the largest count is 8 songs. (b) Mode 2 songs, median 3 songs, mean 3.08 songs. (c) The tallest stack is the mode, the most frequent value, not the center; here the mode is 2 while the median is 3 and the mean is 3.08. (d) Bar heights 11, 12, and 2. The tallest bar now covers 3 to 6 songs even though the most common single value, 2 songs, sits in the first bar, and the gap at 7 songs vanishes.
Problem 2
A clinic records how many minutes past the scheduled time each of 30 patients was seen: 3, 8, 11, 11, 12, 12, 13, 14, 16, 17, 18, 21, 23, 26, 26, 26, 27, 27, 28, 28, 29, 29, 29, 31, 32, 33, 34, 37, 42, 48. (a) Give the bar heights for a histogram with bins of width 10 starting at 0, and describe the shape. (b) Do it again with bins of width 5, and describe the shape. (c) The two descriptions disagree. Which one is wrong? (d) Find the median and say which bin holds it under each width.
Show the worked solution
(a) Bin at width 10, each bin including its left endpoint. From 0 to 10: 3 and 8, so 2 patients. From 10 to 20: 11, 11, 12, 12, 13, 14, 16, 17, 18, so 9. From 20 to 30: 21, 23, 26, 26, 26, 27, 27, 28, 28, 29, 29, 29, so 12. From 30 to 40: 31, 32, 33, 34, 37, so 5. From 40 to 50: 42 and 48, so 2. Check: .
(a) Shape at width 10. The heights 2, 9, 12, 5, 2 rise to one peak in the 20-to-30 bin and fall away on both sides. From this picture you would call the distribution unimodal and roughly symmetric about that bin, with no strong tail either way: 11 of the 30 waits fall below the peak bin and 7 above it.
(b) Bin at width 5. From 0 to 5: 1. From 5 to 10: 1. From 10 to 15: 11, 11, 12, 12, 13, 14, so 6. From 15 to 20: 16, 17, 18, so 3. From 20 to 25: 21, 23, so 2. From 25 to 30: 26, 26, 26, 27, 27, 28, 28, 29, 29, 29, so 10. From 30 to 35: 31, 32, 33, 34, so 4. From 35 to 40: 1. From 40 to 45: 1. From 45 to 50: 1. Check: .
(b) Shape at width 5. The heights 1, 1, 6, 3, 2, 10, 4, 1, 1, 1 show two peaks, one in the 10-to-15 bin and a taller one in the 25-to-30 bin, separated by a dip: only 2 of the 30 patients waited between 20 and 25 minutes. From this picture the distribution is bimodal.
(c) Neither description is wrong. Both are accurate readings of the same 30 numbers. The wide histogram is right that most waits land between 10 and 30 minutes; the narrow one is right that there are two clumps inside that stretch, because the 12 values the wide bar merges are 21, 23, and then ten values bunched from 26 to 29. What is unsupported is the sentence 'the distribution is unimodal', asserted from a single bin width. The shape a histogram shows is a joint property of the data and the width you chose.
(c) The practical rule. Try at least two widths before naming a shape. A common starting point is the range divided by the number of bins you want: the range here is minutes, so about eight bins suggests , rounded to a readable width of 5.
(d) Median. With the median is the average of the 15th and 16th values in the sorted list. Positions 14, 15, and 16 are all 26, so the median is minutes. At width 10 that falls in the 20-to-30 bin, and at width 5 it falls in the 25-to-30 bin. Both times it lands in the tallest bar, which is a coincidence of this data set rather than a rule: the tallest bar is the modal class, and the median need not be inside it.
(a) Heights 2, 9, 12, 5, 2: unimodal with a single peak in the 20-to-30 minute bin. (b) Heights 1, 1, 6, 3, 2, 10, 4, 1, 1, 1: bimodal, with peaks in the 10-to-15 and 25-to-30 bins and only 2 patients between 20 and 25 minutes. (c) Neither. Both read the same 30 values correctly; the unsupported step is naming a shape from one bin width. (d) The median is 26 minutes, in the 20-to-30 bin at width 10 and the 25-to-30 bin at width 5.
Problem 3
A teacher draws a histogram of 20 unit-test scores using bins of width 10 running 60 to 70, 70 to 80, 80 to 90, and 90 to 100. The scores are 62, 65, 66, 68, 70, 70, 73, 76, 79, 80, 80, 80, 83, 85, 88, 90, 92, 95, 97, 99. (a) Give the bar heights when each bin includes its left endpoint and excludes its right. (b) Give them when each bin excludes its left endpoint and includes its right. (c) How many scores moved, and which ones? (d) The teacher says the modal class is the 80s. Is that safe to say from the picture?
Show the worked solution
(a) Left endpoint included. The 60-to-70 bin takes 62, 65, 66, 68, so 4 scores. The 70-to-80 bin takes 70, 70, 73, 76, 79, so 5. The 80-to-90 bin takes 80, 80, 80, 83, 85, 88, so 6. The 90-to-100 bin takes 90, 92, 95, 97, 99, so 5. Check: .
(a) Read the picture. The heights 4, 5, 6, 5 rise to a single tallest bar in the 80s and are close to symmetric around it.
(b) Right endpoint included. Now 70 belongs to the 60-to-70 bin, 80 belongs to the 70-to-80 bin, and 90 belongs to the 80-to-90 bin. The 60-to-70 bin takes 62, 65, 66, 68, 70, 70, so 6. The 70-to-80 bin takes 73, 76, 79, 80, 80, 80, so 6. The 80-to-90 bin takes 83, 85, 88, 90, so 4. The 90-to-100 bin takes 92, 95, 97, 99, so 4. Check: .
(b) Read the picture again. The heights 6, 6, 4, 4 no longer peak in the middle. They start high and step down, which a reader would describe as skewed toward the higher scores rather than symmetric.
(c) Which scores moved. Only the ones sitting exactly on a bin edge: 70, 70, 80, 80, 80, and 90. That is 6 of 20 scores, or 30% of the data, and each one drops exactly one bin when the rule flips. The other 14 scores are strictly inside a bin under both rules and never move.
(d) Not safe from the picture alone. Under the left-endpoint rule the modal class is 80 to 90 with 6 scores; under the right-endpoint rule the 80-to-90 bin holds 4 and is tied for the shortest, while the modal class becomes a tie between 60 to 70 and 70 to 80. Nothing drawn on a histogram records which rule was used, so either the display states its boundary rule in words or a reader needs the raw data to settle the question.
(d) The general point. Boundary effects bite hardest when many values land exactly on the edges, which is common for scores, ages, and anything rounded to a whole number. If you are choosing the edges yourself, put them where no data value can sit, for example at 59.5, 69.5, 79.5, 89.5, and 99.5, and the ambiguity disappears.
(a) 4, 5, 6, 5. (b) 6, 6, 4, 4. (c) Six scores moved, 70, 70, 80, 80, 80, and 90, which is 30% of the data; each drops one bin, and the other 14 never move. (d) No. The modal class is the 80s under the left-endpoint rule but a tie between the 60s and the 70s under the right-endpoint rule, and the drawn histogram does not record which rule was used.
Problem 4
Eighteen fifth-graders each try the same sliding-tile puzzle. The number of seconds each one took is 62, 78, 60, 83, 56, 76, 63, 59, 80, 63, 77, 58, 82, 64, 75, 61, 79, 78. (a) Build a stemplot with stems by tens, and write the key. (b) Describe the shape from that plot. (c) Rebuild it with split stems, low leaves 0 to 4 and high leaves 5 to 9, and describe the shape again. (d) Find the median, and say what part (c) adds to the description.
Show the worked solution
(a) Choose the stems. The times run from 56 to 83 seconds, so the tens digits 5, 6, 7, and 8 give four rows and the ones digit is the leaf. Deal each value onto its row, then order the leaves within each row.
Stem Leaves 5 6 8 9 6 0 1 2 3 3 4 7 5 6 7 8 8 9 8 0 2 3 Key: 5 | 6 represents 56 seconds. Check the leaf count: , which matches the 18 children. Note the repeated leaves, since 63 and 78 each occurred twice.
(b) Shape from the unsplit plot, and the limit of that reading. The row lengths are 3, 6, 6, 3, no stem is empty, and no leaf stands alone, so the outline of the rows is a broad middle falling away evenly at both ends. A reader who looks only at row lengths would call that roughly symmetric and unimodal. Read the leaves and the description does not survive: stem 6 stops at 4 and stem 7 starts at 5, so nothing at all sits between 64 and 75 seconds, a jump of seconds in a range of . That is a gap, and two rows that touch are wide enough to hide it inside an outline. The honest description of these 18 times is two clusters separated by an empty stretch, which is closer to bimodal than to unimodal.
(c) Split the stems. Each stem is written twice, low leaves 0 through 4 on the first copy and high leaves 5 through 9 on the second, which cuts each row from 10 seconds wide to 5. The smallest time is 56 and the largest is 83, so the plot runs from the high copy of stem 5 to the low copy of stem 8.
Stem Leaves 5 (high) 6 8 9 6 (low) 0 1 2 3 3 4 6 (high) 7 (low) 7 (high) 5 6 7 8 8 9 8 (low) 0 2 3 The leaf count is still .
(c) Shape from the split plot. The row lengths are now 3, 6, 0, 0, 6, 3. Two empty rows cover 65 through 74 seconds, so the distribution is bimodal: one cluster of nine children from 56 to 64 seconds and another cluster of nine from 75 to 83 seconds, with nobody at all between them. The unsplit plot let the gap hide because 64 and 75 sat on rows that touch, so its outline looked continuous across the middle even though its leaves did not.
(d) Median. The leaves are already ordered, so counting is the whole job. With the median is the average of the 9th and 10th values. Stems 5 and 6 hold values, so the 9th value is the last leaf on stem 6, which is 64 seconds, and the 10th is the first leaf on stem 7, which is 75 seconds. The median is seconds.
(d) What the split plot adds. The median of 69.5 seconds lands inside the gap: not one of the 18 children took between 65 and 74 seconds. A single center therefore describes nobody in this group, and the honest summary reports the two clusters and their sizes instead. The unsplit plot's leaves carried that gap all along; splitting the stems moved it into the row outline, where no reader can miss it. The row width changed, not the data.
(a) Stem 5 with leaves 6 8 9, stem 6 with 0 1 2 3 3 4, stem 7 with 5 6 7 8 8 9, and stem 8 with 0 2 3, under the key 5 | 6 represents 56 seconds. (b) The row outline, 3, 6, 6, 3, looks roughly symmetric and single-peaked, but the leaves already show nothing between 64 and 75 seconds, a gap of 11 in a range of 27, so the honest reading is two clusters rather than one unimodal group with no gaps. (c) Split rows of 3, 6, 0, 0, 6, 3, so bimodal: nine children from 56 to 64 seconds, nine from 75 to 83 seconds, and an empty stretch from 65 to 74. (d) The median is seconds, which lands in the gap, so no single center describes these children.
Problem 5
A weather station's 20 daily high temperatures, in degrees Fahrenheit, are displayed twice below, first as a stemplot and then as a histogram with bins of width 10.
| Stem | Leaves |
|---|---|
| 6 | 3 7 |
| 7 | 1 2 4 5 6 8 9 |
| 8 | 0 1 3 4 6 7 8 |
| 9 | 0 2 4 5 |
Key: 6 | 3 represents 63 degrees.
| Bin (degrees) | 60 to 70 | 70 to 80 | 80 to 90 | 90 to 100 |
|---|---|---|---|---|
| Days | 2 | 7 | 7 | 4 |
Answer each of these from each display, and where the histogram can only bound the answer, give the bound: (a) the median, (b) the IQR, (c) the number of days at or above 85 degrees, (d) the shape.
Show the worked solution
First check that the two displays hold the same data. The leaf counts by stem are 2, 7, 7, and 4, which are exactly the four bar heights, and both total 20 days. Reading the stemplot back out gives 63, 67, 71, 72, 74, 75, 76, 78, 79, 80, 81, 83, 84, 86, 87, 88, 90, 92, 94, 95.
(a) Median from the stemplot. With the median is the average of the 10th and 11th values. Stem 6 holds 2 values and stem 7 holds 7 more, so 9 values come before stem 8, making the 10th value 80 and the 11th value 81. The median is degrees.
(a) Median from the histogram. The first two bars hold days, so positions 10 through 16 sit in the 80-to-90 bin, and both the 10th and 11th values are in it. The histogram supports only 'the median is between 80 and 90 degrees' and cannot sharpen that, because the seven values in the bar have been replaced by the number 7.
(b) IQR from the stemplot. The lower 10 values are 63, 67, 71, 72, 74, 75, 76, 78, 79, 80, so degrees. The upper 10 are 81, 83, 84, 86, 87, 88, 90, 92, 94, 95, so degrees. Then degrees.
(b) IQR from the histogram. is the average of the 5th and 6th values, and positions 3 through 9 lie in the 70-to-80 bin, so . is the average of the 15th and 16th values, and positions 10 through 16 lie in the 80-to-90 bin, so . Subtracting the extremes, the IQR is somewhere above and below degrees. A bound of 0 to 20 is too wide to report, which is the honest answer here.
(c) Days at or above 85 from the stemplot. On stem 8 the leaves 6, 7, and 8 give 86, 87, 88, and all four values on stem 9 qualify, so 7 days reached 85 degrees or higher.
(c) Days at or above 85 from the histogram. All 4 days in the 90-to-100 bar qualify. The 7 days in the 80-to-90 bar are somewhere from 80 up to but not including 90, and anywhere from 0 to 7 of them could be at or above 85, so the count is between 4 and 11 days. The threshold falls inside a bin, and a histogram cannot split a bin.
(d) Shape from both. The row lengths 2, 7, 7, 4 and the bar heights 2, 7, 7, 4 are the same four numbers, so both displays show the same picture: a broad middle holding 14 of the 20 days across the 70s and 80s, thinner at both ends, with slightly more days in the 90s than in the 60s and no gap or isolated value anywhere. As a check on the near symmetry, the mean is degrees against a median of 80.5 degrees. Shape is the one question here the histogram answers as well as the stemplot.
What the trade buys. The stemplot won every exact question because it kept all 20 values, and it stays readable only because is 20. At 2000 days the leaves would run off the page while four bar heights would still be legible, and that is the point at which you accept the loss and bin.
(a) Stemplot: exactly 80.5 degrees. Histogram: only that the median lies between 80 and 90. (b) Stemplot: , , so degrees. Histogram: only , too wide to be useful. (c) Stemplot: 7 days. Histogram: between 4 and 11 days, because 85 falls inside a bin. (d) Both agree, since the row lengths and bar heights are the same 2, 7, 7, 4: a broad middle across the 70s and 80s, thinner at both ends, no gaps, and roughly symmetric (mean 80.75 against median 80.5).
Problem 6
Two coffee shops each record how long every customer stayed, in minutes, on the same Saturday. Shop A served 200 customers and Shop B served 50. Both histograms use bins of width 15, each including its left endpoint and excluding its right.
| Stay (min) | 0 to 15 | 15 to 30 | 30 to 45 | 45 to 60 | 60 to 75 |
|---|---|---|---|---|---|
| Shop A customers | 20 | 60 | 70 | 40 | 10 |
| Shop B customers | 10 | 20 | 12 | 6 | 2 |
(a) A blogger writes that Shop A's customers linger longer, because every one of A's bars is taller than B's. Evaluate that reasoning. (b) Convert both to relative frequency histograms. (c) Which bin holds each shop's median stay? (d) Is the claim 'more customers stayed 30 to 45 minutes at Shop A than at Shop B' true?
Show the worked solution
(a) The reasoning is invalid, whatever the conclusion turns out to be. Shop A served 200 customers against Shop B's 50, four times as many, so equal shares in a bin would already make A's bar four times taller, and a taller bar therefore says nothing about how long anybody stayed. The totals do not even force what the blogger saw, which holds for these two tables by accident: had B's 50 customers all left within 15 minutes while A's 200 ran 0, 50, 50, 50, 50, A would still have served four times as many people and its first bar would be 0 against B's 50. Comparing raw counts across groups of different sizes measures the group sizes. Frequency histograms compare two groups fairly only when the two totals are equal, and here they are not.
(b) Divide each count by its own shop's total. Shop A over 200: , , , , . These sum to 1.00, which is the check worth doing.
(b) Shop B over 50: , , , , . These also sum to 1.00. Now the two histograms sit on the same vertical scale and the bars are comparable.
(b) Read the shares. In the two shortest bins Shop B leads: 20% of B's customers left within 15 minutes against 10% of A's, and 40% of B's stayed 15 to 30 minutes against 30% of A's. In the three longest bins Shop A leads, 35% against 24%, 20% against 12%, and 5% against 4%. So A's customers did stay longer, but that conclusion comes from the shares, not from the taller bars the blogger pointed at.
(c) Shop A's median. With it is the average of the 100th and 101st stays. The cumulative counts through the bins are 20, 80, 150, 190, 200, so positions 81 through 150 lie in the 30-to-45 bin and both the 100th and the 101st do. A's median stay is between 30 and 45 minutes.
(c) Shop B's median. With it is the average of the 25th and 26th stays. The cumulative counts are 10, 30, 42, 48, 50, so positions 11 through 30 lie in the 15-to-30 bin and both the 25th and the 26th do. B's median stay is between 15 and 30 minutes. The two medians land in different bins, one bin apart, which agrees with what the shares showed.
(d) The claim is true as written, and it is a different claim from the blogger's. In raw counts, 70 of Shop A's customers stayed 30 to 45 minutes against 12 of Shop B's, so more people did that at A. As a share of each shop's own customers it was 35% at A against 24% at B, which is also true and much closer. Both sentences are correct because they answer different questions, so say which one you mean: more people, or a larger share of that shop's people.
(a) Invalid, even though the conclusion survives. A served four times as many customers, so equal shares alone would make its bars four times taller, and taller bars are not even forced by the totals: B's 50 could all have landed in one bin where A had none. Counts across unequal groups compare group sizes. (b) A: 0.10, 0.30, 0.35, 0.20, 0.05. B: 0.20, 0.40, 0.24, 0.12, 0.04. Each set sums to 1.00. (c) A's median is in the 30-to-45 bin (100th and 101st of 200) and B's is in the 15-to-30 bin (25th and 26th of 50). (d) True in raw counts, 70 against 12, though the shares are 35% against 24%.
Problem 7
A nursery grows seedlings in two trays and measures the height, in centimeters, of 11 seedlings from each.
- Tray 1: 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30
- Tray 2: 10, 12, 14, 14, 15, 20, 25, 26, 26, 28, 30
The grower draws side-by-side boxplots and concludes the two trays are essentially identical. (a) Give both five-number summaries. (b) Give both means and both sample standard deviations. (c) Is the grower's conclusion supported by the boxplots? (d) Draw both as dotplots instead and say what changes.
Show the worked solution
(a) Both lists are already sorted and each has , so each median is the 6th value and, the count being odd, the median is left out of both halves. Tray 1: median 20 cm, lower half 10, 12, 14, 16, 18 giving , upper half 22, 24, 26, 28, 30 giving , minimum 10, maximum 30. The summary is 10, 14, 20, 26, 30 cm.
(a) Tray 2: median 20 cm, lower half 10, 12, 14, 14, 15 giving , upper half 25, 26, 26, 28, 30 giving , minimum 10, maximum 30. The summary is 10, 14, 20, 26, 30 cm, identical to Tray 1's.
(a) So the two boxplots match line for line: the same box from 14 to 26 cm, the same median line at 20 cm, and the same whiskers to 10 and 30 cm. Both have cm, and with the fences are cm and cm, so neither tray has an outlier and the modified boxplot is the same drawing as the basic one for both.
(b) Means. Tray 1 sums to and Tray 2 sums to , so both means are cm, equal to both medians.
(b) Standard deviations. Tray 1's squared deviations from 20 are 100, 64, 36, 16, 4, 0, 4, 16, 36, 64, 100, summing to 440, so cm. Tray 2's are 100, 64, 36, 36, 25, 0, 25, 36, 36, 64, 100, summing to 522, so cm. The standard deviation is the first summary number that separates the two trays, and a boxplot never shows it.
(c) The conclusion is not supported. Identical boxplots establish exactly one thing: the two trays match on those five numbers. A boxplot draws the region from to as a solid rectangle, so it looks the same whether the values inside are spread evenly or piled against the two ends, and it carries no mean, no standard deviation, and no sample size. 'The boxplots are identical' is a true statement about the display, not about the trays.
(d) Dotplots of the same numbers. Tray 1 puts one dot above every even height from 10 to 30 cm, a flat row of eleven single dots with no peak at all, which is close to a uniform shape. Tray 2 puts five dots between 10 and 15 cm, one lone dot at 20 cm, and five dots between 25 and 30 cm, so it is clearly bimodal with a thinly occupied middle: only one of its eleven seedlings measured between 16 and 24 cm, against five of Tray 1's.
(d) What changed. Nothing about the data; only the display. A boxplot is a picture of five summary statistics, so it can carry only what those five numbers carry, and clusters, gaps, extra peaks, and are not among them. Tray 2's split is the kind of thing worth chasing down, plausibly two seed lots or two positions under the lights, and the boxplot gave the grower no reason to look.
(a) Both are 10, 14, 20, 26, 30 cm, so the two boxplots are identical, with cm and no outliers in either tray. (b) Both means are 20 cm; the sample standard deviations differ, 6.63 cm for Tray 1 against 7.22 cm for Tray 2. (c) No. Identical boxplots show only that those five numbers match, and the box is drawn solid whatever sits inside it. (d) The dotplots differ sharply: Tray 1 is a flat row of single dots at every even height, close to uniform, while Tray 2 is bimodal, with clusters at 10 to 15 cm and 25 to 30 cm and one seedling at 20 cm between them.
Problem 8
A food co-op weighs every bulk-bin purchase made in one afternoon and draws a histogram of the 80 weights, in pounds, with bins of width 5, each including its left endpoint and excluding its right.
| Weight (lb) | 0 to 5 | 5 to 10 | 10 to 15 | 15 to 20 | 20 to 25 | 25 to 30 |
|---|---|---|---|---|---|---|
| Purchases | 6 | 16 | 22 | 20 | 10 | 6 |
(a) Confirm the total and give the percent of purchases under 15 pounds. (b) Which bin contains the median? (c) Which bins contain and ? (d) What can you say about the IQR? (e) The manager reports the median as 12.5 pounds, the middle of the tallest bin. Is that right?
Show the worked solution
(a) Total: purchases, which matches the stated count. Under 15 pounds means the first three bins, purchases, and , so 55% weighed under 15 pounds.
(a) Build the cumulative counts once and reuse them for every part: 6 through the first bin, 22 through the second, 44 through the third, 64 through the fourth, 74 through the fifth, and 80 through the sixth.
(b) Median. With the median is the average of the 40th and 41st values. Positions 23 through 44 sit in the 10-to-15 bin, so both the 40th and the 41st do, and the median is between 10 and 15 pounds.
(c) . The count is even, so the median falls between two values and the lower half is the first 40 purchases. is the median of those, the average of the 20th and 21st values overall. Positions 7 through 22 sit in the 5-to-10 bin, so both do, and is between 5 and 10 pounds.
(c) is the median of the upper 40, the average of the 60th and 61st values overall. Positions 45 through 64 sit in the 15-to-20 bin, so both do, and is between 15 and 20 pounds.
(d) IQR. Combining with , the difference is smallest when is at 15 and is just under 10, giving a value just above 5, and largest when is just under 20 and is 5, giving a value just under 15. So all the histogram supports is pounds. If the IQR matters, get the data or a display that kept it.
(e) Half right, and wrong for the reason that counts. The median does lie between 10 and 15 pounds, so 12.5 is not contradicted by anything here. But 12.5 is the midpoint of a bin, which is a fact about where the co-op drew its bin edges, not a number read from any purchase: the 22 values in that bin could all be near 10.1 pounds or all near 14.9. Redraw the histogram with a width of 10 and the same median sits in a 10-to-20 bin whose midpoint is 15. The histogram supports 'the median is between 10 and 15 pounds' and nothing sharper.
(a) 80 purchases in total, and 44 of them, or 55%, weighed under 15 pounds. (b) The 10-to-15 bin, which holds the 40th and 41st values. (c) is in the 5-to-10 bin (20th and 21st values) and is in the 15-to-20 bin (60th and 61st). (d) Only that pounds. (e) The median does fall between 10 and 15 pounds, so 12.5 is not contradicted, but a bin midpoint is a property of where the edges were drawn rather than a value from the data, and the histogram supports only the interval.