What is a bimodal distribution? Two peaks, explained

By Jude Wallis · Updated

A bimodal distribution has two prominent peaks. Two peaks almost always mean two groups have been mixed into one data set, so the fix is to split the data and describe each group rather than summarize everything with one mean. The value between the peaks is often the rarest one.

AP Statistics: Unit 1 (topics 1.6 Descriptions for One Quantitative Variable Distributions, 1.5 Graphical Representations for One Quantitative Variable, 1.9 Comparisons of the Distributions for One Quantitative Variable). Bimodal is defined in Unit 1 topic 1.6 of the Fall 2026 AP Statistics course (1.6.A.3), alongside the gaps and clusters you report with it (1.6.A.5 and 1.6.A.6); topic 1.9 is where the framework lists shape, clusters, and gaps as readable from histograms, stemplots, and dotplots but not from boxplots.

What makes a distribution bimodal

A distribution is bimodal when its graph shows two prominent peaks. The Fall 2026 AP course framework sets out the vocabulary in topic 1.6: one main peak makes a distribution unimodal, two prominent peaks make it bimodal, and a shape whose frequencies are all about equal, with no prominent peak, is approximately uniform.

The word doing the work is prominent. Real data are bumpy, so nearly every histogram has small wiggles that are not peaks. Look for two clear high regions separated by a real dip, deep enough that you would naturally call the two humps separate.

Bimodal is not a kind of skew. Skew describes one tail stretching further than the other away from a single peak, so a shape with two peaks needs its own label. Say bimodal first, then say roughly where the two peaks sit. If you are still shaky on skew, see skewed left vs skewed right.

Two peaks usually mean two groups

Bimodality is a clue about where the data came from, not just a description of a picture. When you see two peaks, the first thing to suspect is that two different groups were poured into one data set. Each group has its own center, and the graph shows both of them.

The usual suspects are easy to name once you look for them.

  • Two sections of the same class, where one section had a review session and the other did not.
  • Two machines filling the same product, where one drifted out of calibration.
  • Adults and children measured together, so heights or reaction times pile up in two places.
  • Two time periods sharing one graph, such as lunch orders and dinner orders at a restaurant.

The repair follows from the diagnosis. Split the data by the grouping variable and describe each group on its own, because a summary of the combined data describes a population that does not exist. When you cannot identify the grouping variable, say so and report the two peaks anyway, since naming the pattern is more honest than averaging it away.

Why one mean and one standard deviation describe it badly

The mean is the balance point of the data, so when there are two clumps the balance point lands between them. The standard deviation, written ss for a sample and σ\sigma (sigma) for a population, measures typical distance from the mean. With two clumps, most of that distance comes from the gap between the groups rather than from variability inside either group.

Here is the effect on real numbers. Twelve students take a 20-point quiz, six from a section that had a review session and six from a section that did not: 6, 7, 7, 8, 8, 9, 16, 17, 17, 18, 18, 19.

Taken as one group, the mean is xˉ\bar{x} (read x-bar) =12.5= 12.5 points with s5.32s \approx 5.32 points. Split into the two sections, the low section has a mean of 7.5 with s1.05s \approx 1.05, and the high section has a mean of 17.5 with s1.05s \approx 1.05. The combined standard deviation is about five times larger than either group's, and it is measuring the distance between the sections, not how spread out students are within a section.

One more consequence. Rules built for a single-peaked, symmetric shape do not apply here, so do not reach for the empirical rule or a normal-model calculation on bimodal data.

The value between the peaks can be the least typical value

This is the part that surprises students. In the quiz data above, the mean is 12.5 and the median is also 12.5, yet nobody scored 12.5, and in fact nobody scored anywhere between 9 and 16. The center of the distribution sits inside an empty stretch.

ScoreCount
61
72
82
91
10 to 150
161
172
182
191

So the sentence "students typically scored about 12.5" describes no student in the room. Central and typical travel together for a unimodal distribution, and they come apart for a bimodal one. That is the whole argument for splitting the data: 7.5 and 17.5 are values students actually scored near, and 12.5 is not.

The AP framework asks you to report unusual features such as outliers, gaps, and clusters in context. A bimodal distribution normally hands you two of those at once: two clusters of values with a gap between them.

Histograms and dotplots show it; bin width can hide it

A dotplot places one dot for each value, so it shows every gap and cluster exactly as they occur. For a data set small enough to plot dot by dot, this is the most honest picture of modality you can draw, and a stem-and-leaf plot does much the same job.

A histogram groups values into bins first, and the framework notes that changing the bin widths changes the appearance of the histogram. That cuts both ways. Bins that are too wide can swallow the dip between two peaks and make a bimodal data set look unimodal, while bins that are too narrow can turn ordinary sampling noise into a row of fake peaks.

The practical habit is to redraw the histogram with a second bin width before you commit to a shape. If the two peaks survive a reasonable change in bin width, they are real. For the trade-offs between these graphs, see dotplot vs histogram vs stemplot.

Why a boxplot hides bimodality completely

A boxplot draws exactly five numbers: the minimum, the first quartile, the median, the third quartile, and the maximum. None of those five counts peaks. The box tells you where the middle 50% of the data lies, but nothing about how the values are arranged inside the box, so a boxplot cannot show modality at all. (This site uses the median-excluded rule for quartiles, the one built into the TI-84.)

The framework's own lists draw the same line. Where it describes what you can compare using histograms, stem-and-leaf plots, and dotplots, shape and clusters and gaps are included. Where it describes what you can compare using boxplots, the list stops at center, variability, outliers, and skew or symmetry. The features left off the second list are exactly the ones bimodality lives in.

Run the quiz data through it. The five-number summary is 6, 7.5, 12.5, 17.5, 19, giving a box from 7.5 to 17.5 with the median dead center at 12.5 and two whiskers of equal length. The boxplot looks textbook symmetric with no outliers, and it gives no hint that the middle of that box is empty.

So if a question shows you a boxplot and asks about shape, the safe answer names symmetry or skew and stops there. To build one from scratch, see how to make a boxplot.

How to describe a bimodal distribution on the exam

A complete description of one quantitative variable covers shape, center, and variability, plus unusual features, always in context. Bimodality changes what you write in every one of those slots.

  1. Shape. Say bimodal, and give the approximate location of each peak.
  2. Center. Report the median if you must give one number, then say plainly that a single center is misleading here because the data fall in two clusters.
  3. Variability. Give the range or the interquartile range, and note that the spread mostly reflects the distance between the two clusters.
  4. Unusual features. Name the gap between the clusters and the clusters themselves.
  5. Context. If the setting suggests two groups, say what they probably are and recommend describing them separately.

A one-sentence version reads like this: the scores are bimodal with clusters near 6 to 9 points and near 16 to 19 points and a gap between 9 and 16, so a single mean of 12.5 points describes no student and the two sections should be summarized separately. You can practice writing descriptions like that in the describing distributions problem set.

Common mistakes

  • Calling a distribution bimodal because of any two bumps. The peaks have to be prominent, with a real dip between them.
  • Reporting a mean and standard deviation as if they summarize the data. They are still computable, but they describe a single group that is not there.
  • Judging modality from a boxplot. A boxplot cannot show peaks, so a bimodal set and an evenly spread set can produce identical boxes.
  • Trusting one histogram. Change the bin width once before you decide.
  • Calling it skewed. Skew describes one tail relative to a single peak, so a two-peaked shape is described as bimodal instead.
  • Applying the empirical rule or a normal calculation. Those need a single-peaked, roughly symmetric shape.

Splitting a bimodal data set fixes the summary

Twelve students take a 20-point quiz. Six are from a section that attended a review session and six are from a section that did not, but the scores were recorded together: 6, 7, 7, 8, 8, 9, 16, 17, 17, 18, 18, 19. Find the mean and sample standard deviation for the combined data, then for each section, and say which summary you would report.

  1. Combined sum: 6+7+7+8+8+9+16+17+17+18+18+19=1506 + 7 + 7 + 8 + 8 + 9 + 16 + 17 + 17 + 18 + 18 + 19 = 150, with n=12n = 12 values.

  2. Combined mean: xˉ=15012=12.5\bar{x} = \frac{150}{12} = 12.5 points.

  3. Squared deviations from 12.5 for the low six: (6.5)2=42.25(-6.5)^2 = 42.25, (5.5)2=30.25(-5.5)^2 = 30.25, (5.5)2=30.25(-5.5)^2 = 30.25, (4.5)2=20.25(-4.5)^2 = 20.25, (4.5)2=20.25(-4.5)^2 = 20.25, (3.5)2=12.25(-3.5)^2 = 12.25. These add to 155.5.

  4. The high six are the mirror image (+3.5+3.5 through +6.5+6.5), so they add to 155.5 as well. Total sum of squares =155.5+155.5=311= 155.5 + 155.5 = 311.

  5. Combined sample standard deviation: s=311121=31111=28.27275.32s = \sqrt{\frac{311}{12 - 1}} = \sqrt{\frac{311}{11}} = \sqrt{28.2727} \approx 5.32 points.

  6. Now split. Low section: 6+7+7+8+8+9=456 + 7 + 7 + 8 + 8 + 9 = 45, so its mean is 456=7.5\frac{45}{6} = 7.5 points.

  7. Low section squared deviations from 7.5: 2.25,0.25,0.25,0.25,0.25,2.252.25, 0.25, 0.25, 0.25, 0.25, 2.25, which add to 5.5. So s=5.561=1.11.05s = \sqrt{\frac{5.5}{6 - 1}} = \sqrt{1.1} \approx 1.05 points.

  8. High section: 16+17+17+18+18+19=10516 + 17 + 17 + 18 + 18 + 19 = 105, so its mean is 1056=17.5\frac{105}{6} = 17.5 points, and by the same mirror-image arithmetic s1.05s \approx 1.05 points.

  9. Compare: one summary says 12.5 with a spread of about 5.32; two summaries say 7.5 and 17.5, each with a spread of about 1.05.

Combined: mean 12.5 points, s5.32s \approx 5.32 points. Split: the low section has mean 7.5 with s1.05s \approx 1.05, and the high section has mean 17.5 with s1.05s \approx 1.05. Report the two sections separately. The combined mean of 12.5 matches no student, and most of the combined standard deviation is measuring the distance between sections rather than variation within either one.

Two very different data sets, one identical boxplot

Set A (the bimodal quiz scores) is 6, 7, 7, 8, 8, 9, 16, 17, 17, 18, 18, 19. Set B is 6, 7, 7, 8, 9, 12, 13, 15, 17, 18, 18, 19. Find the five-number summary of each, using the median-excluded (TI-84) quartile rule, and explain what that shows about boxplots.

  1. Both sets are already sorted and both have n=12n = 12, an even count, so the lower half is the first 6 values and the upper half is the last 6.

  2. Set A median: the 6th and 7th values are 9 and 16, so the median is 9+162=12.5\frac{9 + 16}{2} = 12.5.

  3. Set A lower half 6, 7, 7, 8, 8, 9: its middle two values are 7 and 8, so Q1=7+82=7.5Q_1 = \frac{7 + 8}{2} = 7.5.

  4. Set A upper half 16, 17, 17, 18, 18, 19: its middle two values are 17 and 18, so Q3=17+182=17.5Q_3 = \frac{17 + 18}{2} = 17.5. Minimum 6, maximum 19.

  5. Set B median: the 6th and 7th values are 12 and 13, so the median is 12+132=12.5\frac{12 + 13}{2} = 12.5.

  6. Set B lower half 6, 7, 7, 8, 9, 12: middle two are 7 and 8, so Q1=7+82=7.5Q_1 = \frac{7 + 8}{2} = 7.5. Upper half 13, 15, 17, 18, 18, 19: middle two are 17 and 18, so Q3=17+182=17.5Q_3 = \frac{17 + 18}{2} = 17.5. Minimum 6, maximum 19.

  7. Both five-number summaries are 6, 7.5, 12.5, 17.5, 19, so IQR=17.57.5=10IQR = 17.5 - 7.5 = 10 for both and neither set has an outlier, since 1.5×10=151.5 \times 10 = 15 puts the fences at 7.515=7.57.5 - 15 = -7.5 and 17.5+15=32.517.5 + 15 = 32.5.

  8. The histograms are nothing alike. Set A has nothing at all between 9 and 16, while Set B has three values (12, 13, 15) filling that stretch, and Set B's values step up steadily across the range instead of piling into two clumps.

Both sets share the five-number summary 6, 7.5, 12.5, 17.5, 19, so they produce identical boxplots, yet Set A is bimodal with an empty gap from 9 to 16 and Set B has values spread across the whole range. A boxplot cannot distinguish them, which is why you need a histogram or dotplot before describing shape.

Frequently asked questions

Is a bimodal distribution the same as a distribution with two modes?

Not quite. The mode is the most frequent value, and a list can have several, including two tied values that sit right next to each other, which is not a bimodal shape. Bimodal describes two prominent peaks in the graph, separated by a real dip, so check the picture rather than the tie.

Can a bimodal distribution be symmetric?

Yes. If the two peaks are about the same height and sit at equal distances from the center, the shape is symmetric and bimodal at the same time. Symmetry and modality are separate features, so describe both.

Should I ever report the mean of bimodal data?

Only with a warning attached. The mean is still a correct calculation, but it lands between the two clusters and describes a case that may not occur in the data. Report the two peak locations, or split the data by group, and say why.

How do I tell a bimodal shape from noisy data?

Redraw the graph with a different bin width, or use a dotplot that shows every value. Two peaks that survive a reasonable change in bin width, with a clear gap or dip between them, are real. Two bumps that vanish when you widen the bins were noise.

What causes bimodal data in the first place?

Almost always a mixture of two groups: two class sections, two machines, two time periods, or two populations such as adults and children. Look for a variable in the context that would split the data into the two clusters you can see.