How to describe the shape of a distribution

By Jude Wallis · Published

Name two things: symmetry or skew, and the number of peaks. A distribution is approximately symmetric, skewed left, or skewed right, and unimodal, bimodal, or approximately uniform. Skew is named for the side the long tail points to, never for the side the tall bars are on.

AP Statistics: Unit 1 (topics 1.6 Descriptions for One Quantitative Variable Distributions, 1.9 Comparisons of the Distributions for One Quantitative Variable). Shape vocabulary belongs to Unit 1 of the Fall 2026 AP Statistics course: naming a distribution's symmetry, skew, and number of peaks in topic 1.6, and comparing shapes across distributions in topic 1.9. Unit 1 is the most heavily weighted unit on the multiple-choice section at 20 to 30 percent.

Shape is two words, not one

A full shape description has two parts, and answers that give only one part leave credit on the table.

  1. Symmetry or skew. Approximately symmetric, skewed left, or skewed right.
  2. Number of peaks. Unimodal (one peak), bimodal (two peaks), or approximately uniform (no real peak).

So a complete shape reads like "unimodal and skewed right" or "bimodal and roughly symmetric." Add any unusual features you can see, meaning outliers, gaps, and clusters, and the shape part of your description is done.

The word approximately is doing real work. Actual data are never perfectly symmetric and almost never perfectly uniform, so "approximately symmetric" and "roughly symmetric" are the correct phrasings, and a bare "symmetric" claims more than any real histogram supports.

This page covers the shape vocabulary only. Shape is one quarter of a full description of a distribution; for the center, variability, and the sentence structure that scores well, see how to describe a distribution.

Skew is named for the tail, not the bulk

This is the single most reversed idea in the whole unit, so it deserves its own rule.

A distribution is skewed toward its longer tail. A long tail stretching toward the large values is skewed right, even though the tall bars sit on the left. A long tail stretching toward the small values is skewed left, even though the tall bars sit on the right.

Students get this backward because "the data are mostly on the left" feels like it should be called left-skewed. It is not. The name follows the thin, stretched-out end.

The reason the tail gets naming rights is that the tail is what makes the distribution behave unusually. It is the tail that drags the mean away from the median, that widens the standard deviation, and that decides whether you should report a resistant summary. Two ways to keep it straight:

  • Point at the tail. Physically point along the long tail. Right means skewed right.
  • Use the mean. In a right-skewed distribution the mean is usually greater than the median, because the long right tail pulls it up. Left skew usually puts the mean below the median. Compute both and the gap will usually tell you the direction. This is a strong tendency rather than a theorem, so confirm it with a graph. See mean vs median and skewed left vs skewed right.

Also common: "positively skewed" is the same thing as skewed right, and "negatively skewed" is the same as skewed left.

Counting peaks: unimodal, bimodal, uniform

A peak (or mode, in the shape sense) is a place where the graph is locally tallest, with lower values on both sides.

  • Unimodal: one clear peak. Most distributions you meet in Unit 1 are unimodal.
  • Bimodal: two clear peaks, usually separated by a dip. A bimodal shape often means two different groups got mixed into one graph, for example the heights of adult men and adult women together, or exam scores from a class that split into students who studied and students who did not. When you spot one, saying what the two groups might be is a strong observation.
  • Approximately uniform: every bar is about the same height, with no peak at all. Rolling a fair die many times produces this. See uniform distribution explained.

Two cautions. First, a bar being one pixel taller than its neighbor is not a peak; you are looking for a prominent rise with a real dip beside it. Second, a bimodal graph shape is not the same thing as the data having two modes in the counting sense, where the mode is the most frequently occurring value. A distribution can have exactly one most-common value and still look clearly bimodal, and two values can tie for most common without producing two peaks anywhere. See how to find the mode and what is a bimodal distribution.

A note on graph choice: histograms can hide or invent peaks depending on the bin width. If a shape looks borderline, check it on a dotplot or a stemplot, where every value is visible. Compare the three in dotplot vs histogram vs stemplot.

Gaps, clusters, and outliers

Three features get named alongside shape.

  • An outlier is a value far from the rest of the data. On a graph it is a lone dot or bar separated from everything else.
  • A gap is an interval where no values occur at all. Gaps are properties of the empty space, not of any value.
  • A cluster is a group of values bunched together, usually cut off from other values by a gap.

They travel together. An outlier is almost always on the far side of a gap, and a bimodal shape is usually two clusters with a gap between them.

The difference between "outlier" and "cluster" is a matter of how many values are out there. One lone value past a gap is an outlier. Six values past the same gap are a second cluster, and a second cluster is a much stronger hint that two different kinds of individual are in your data.

On a graph you identify a possible outlier by eye. To back that up with a rule, use the 1.5×IQR1.5 \times \text{IQR} criterion: a value is an outlier if it is more than 1.5×IQR1.5 \times \text{IQR} above Q3Q_3 or below Q1Q_1. Work it in the 1.5 IQR rule. The AP course also allows a second rule, more than 2 standard deviations from the mean.

Shape decides which summaries you report

Naming the shape is not decoration. It picks your numbers.

Shape you seeReport center asReport variability as
Approximately symmetric, no outliersMean xˉ\bar{x}Standard deviation ss
Skewed either directionMedianInterquartile range
Any shape with outliersMedianInterquartile range

The logic is resistance. The median and IQR barely move when a far-out value is added, while the mean and standard deviation chase it. See what is a resistant statistic and standard deviation vs IQR.

Bimodal shapes deserve an extra sentence, because a single measure of center can be actively misleading there. In the second worked example below, ten values split into a low cluster and a high cluster produce a mean of 6.7 and a median of 6.5, and not one observation is anywhere near either number. Reporting "the center is about 6.6" would describe a value the data never produced. When a distribution is bimodal, say so and describe the two clusters separately.

Shape also drives conditions later in the course. Strong skew is the reason a small-sample t-procedure can fail its sample data condition, which is where the normal probability plot becomes useful.

Where shape gets graded

Shape appears in Unit 1 of the Fall 2026 AP Statistics course, in topic 1.6 (describing a quantitative distribution) and again in topic 1.9 (comparing two or more distributions). Unit 1 is the most heavily weighted unit on the multiple-choice section at 20 to 30 percent, so this vocabulary is worth more than its difficulty suggests.

Three things cost points here more than anything else.

  1. Naming shape with no context. "Skewed right" is half an answer. "The distribution of daily absences is skewed right" is the answer.
  2. Reversing the skew direction. Point at the tail.
  3. Stopping after shape. Shape alone is not a description of a distribution. Center, variability, and unusual features all have to appear, each with a number and units where one applies.

When you compare two distributions, name the shape of each one and then compare them explicitly, using words like "more strongly skewed than" rather than describing them in separate paragraphs and leaving the comparison to the reader. See how to compare two distributions, then drill it in describing distributions practice or the descriptive statistics sandbox.

Describe the shape of a dotplot of absences

A teacher records days absent for 15 students: 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 3, 4, 4, 9, 10. Describe the shape of this distribution, and justify the direction of skew with the summary statistics.

  1. Build the dotplot counts: the value 1 appears 3 times, 2 appears 5 times, 3 appears 3 times, 4 appears 2 times, 9 appears once, and 10 appears once. Values 5 through 8 do not appear at all.

  2. Find the peak. The tallest stack is at 2 days with 5 students, and the counts fall off on both sides of it, so the distribution is unimodal.

  3. Find the tail. The bulk of the data sits between 1 and 4 days, and the two values at 9 and 10 stretch far to the right. The long tail points toward the large values, so the distribution is skewed right.

  4. Name the gap. There are no observations at 5, 6, 7, or 8 days, so there is a gap between 4 and 9.

  5. Check for outliers with the 1.5×IQR1.5 \times \text{IQR} rule. Sorted, n=15n = 15, so the median is the 8th value, 2. Excluding the median, the lower half is 1, 1, 1, 2, 2, 2, 2 with median Q1=2Q_1 = 2, and the upper half is 3, 3, 3, 4, 4, 9, 10 with median Q3=4Q_3 = 4. So IQR=42=2\text{IQR} = 4 - 2 = 2.

  6. Compute both fences. The upper fence is Q3+1.5×IQR=4+1.5(2)=4+3=7Q_3 + 1.5 \times \text{IQR} = 4 + 1.5(2) = 4 + 3 = 7, and both 9 and 10 are above it, so both are outliers. The lower fence is Q11.5×IQR=23=1Q_1 - 1.5 \times \text{IQR} = 2 - 3 = -1, and nobody was absent a negative number of days, so there are no low outliers.

  7. Confirm the skew direction with center. The sum is 49 and n=15n = 15, so xˉ=49/15=3.2667\bar{x} = 49/15 = 3.2667 days, while the median is 2 days. The mean sits above the median, which is the usual sign of right skew.

The distribution of days absent is unimodal with a single peak at 2 days, skewed right, with a gap between 4 and 9 days and two high outliers at 9 and 10 days (both above the upper fence of 7). The mean of 3.27 days sits above the median of 2 days, consistent with the right skew.

A bimodal shape where the center describes nobody

Ten commute distances, in miles, are 2, 3, 3, 4, 4, 9, 10, 10, 11, 11. Describe the shape, and explain why reporting a single measure of center would mislead.

  1. Look at where the values sit. Five values fall between 2 and 4 miles, and five fall between 9 and 11 miles. Nothing lands between 5 and 8 miles.

  2. Name the shape. There are two separated groups of values, so the distribution is bimodal, with a low cluster from 2 to 4 miles and a high cluster from 9 to 11 miles, separated by a gap from 5 to 8 miles.

  3. Say whether it is skewed. Neither cluster has a long tail running away from it, and the two clusters are about the same size and spread, so there is no overall direction to the skew: call it roughly symmetric. The mean and median computed in the next two steps differ by only 0.2 miles, far too small a gap to read as skew, and the mean-versus-median rule is not trustworthy for a bimodal shape anyway. Bimodal and roughly symmetric can be true at the same time.

  4. Compute the mean. The sum is 2+3+3+4+4+9+10+10+11+11=672 + 3 + 3 + 4 + 4 + 9 + 10 + 10 + 11 + 11 = 67, so xˉ=67/10=6.7\bar{x} = 67/10 = 6.7 miles.

  5. Compute the median. With n=10n = 10, average the 5th and 6th sorted values: (4+9)/2=6.5(4 + 9)/2 = 6.5 miles.

  6. Compare those to the data. No commute in the data set is between 5 and 8 miles, so neither 6.7 nor 6.5 is close to any observed value.

  7. Check that the outlier rule finds nothing. Q1=3Q_1 = 3 and Q3=10Q_3 = 10, so IQR=7\text{IQR} = 7, and the fences at 31.5(7)=7.53 - 1.5(7) = -7.5 and 10+1.5(7)=20.510 + 1.5(7) = 20.5 leave every value inside. The unusual feature here is the gap, not an outlier.

The distribution of commute distances is bimodal and roughly symmetric, with a cluster from 2 to 4 miles, a cluster from 9 to 11 miles, and a gap from 5 to 8 miles. Both the mean (6.7 miles) and the median (6.5 miles) land inside the gap, where no commute actually falls, so the two clusters should be described separately rather than summarized by one center.

Frequently asked questions

Is skew named for the tail or for the peak?

For the tail. A distribution is skewed toward its longer tail, so a long right tail means skewed right even though the tall bars sit on the left. Point along the thin stretched-out end and the direction you are pointing is the name.

Can a distribution be both bimodal and symmetric?

Yes. Symmetry and the number of peaks are separate questions. Two equally sized clusters with a gap between them, such as 2, 3, 3, 4, 4, 9, 10, 10, 11, 11, is bimodal and roughly symmetric at the same time. Name both parts.

What is the difference between a gap and an outlier?

A gap is an interval containing no data at all; an outlier is a single value sitting far from the rest. An outlier is almost always separated from the bulk of the data by a gap, but a gap can exist with a whole cluster on the far side of it rather than one lone point.

Do I have to say approximately when naming a shape?

Use it for symmetric and uniform. Real data are never exactly symmetric or exactly uniform, so "approximately symmetric" and "approximately uniform" are the accurate phrasings. Skewed right and skewed left are already descriptions of a tendency, so they do not need the hedge.

How does shape change which statistics I report?

Approximately symmetric with no outliers means report the mean and the standard deviation. Skewed in either direction, or any shape with outliers, means report the median and the interquartile range, because those two are resistant. For a bimodal shape, describe the clusters instead of leaning on one center.