How to describe a distribution (and get full credit)

By Jude Wallis · Updated

Describe a distribution with four features in context: shape (skewed right, skewed left, or symmetric), center (a median or mean with units), variability (an IQR or standard deviation with units), and unusual features (outliers, gaps, clusters). Naming a measure without a value earns nothing.

AP Statistics: Unit 1 (topics 1.6 Descriptions for One Quantitative Variable Distributions, 1.7 Summary Statistics for One Quantitative Variable). In the Fall 2026 AP Statistics course, describing a one-variable quantitative distribution by shape, center, variability, and unusual features in context is Unit 1 topic 1.6 (1.6.A.1), scored under skill 4.A.

The four features a full-credit description names

The AP course asks for four things when you describe the distribution of one quantitative variable: shape, center, variability (also called spread), and any unusual features such as outliers, gaps, or clusters. All four have to appear in context, which means naming the variable, the group that was measured, and the units.

The College Board is blunt about where students lose points here. Many write out an acronym like SOCS and then never explain the pieces, and unusual features get skipped most often. A response that covers only the center is scored as incomplete.

So treat the four features as a checklist you write out in full sentences, not a mnemonic you recite. This is topic 1.6 Descriptions for One Quantitative Variable Distributions in Unit 1.

A sentence frame you can reuse

Four features, four sentences. Fill in the blanks and you have a complete answer every time.

  1. Shape. The distribution of [variable] for [the group measured] is [skewed to the right / skewed to the left / roughly symmetric], with [one peak / two peaks].
  2. Center. The [median / mean] [variable] is about [value] [units].
  3. Variability. The [IQR / range / standard deviation] is about [value] [units].
  4. Unusual features. There is [an outlier at (value) (units) / a gap between (a) and (b) / a cluster near (value)], or: there are no apparent outliers, gaps, or clusters.

The frame forces the two things a grader is scanning for: a number with units in every sentence, and the variable named at least once. If a feature genuinely is not present, say so out loud. Writing that there are no apparent outliers is an answer, while saying nothing about outliers is a hole in the response.

Partial credit and full credit on the same histogram

Here is a histogram of the amount each of 40 shoppers spent at one farmers market stall, written as a frequency table.

Amount spent (dollars)Number of shoppers
0 to under 1012
10 to under 2014
20 to under 307
30 to under 404
40 to under 501
50 to under 600
60 to under 702

A typical partial-credit answer: "The distribution is skewed right. The center is the median and the spread is the range. There are a few outliers on the right."

A full-credit answer: "The distribution of amounts spent by these 40 shoppers is skewed to the right, with a single peak in the 10 to 20 dollar bin. The median falls in that same bin (the 20th and 21st values both land there), so a typical shopper spent roughly 15 dollars. Amounts run from 0 to 70 dollars, a range of about 70 dollars, and most of the data is packed below 30 dollars. There is a gap from 50 to 60 dollars, and the 2 shoppers who spent between 60 and 70 dollars stand apart from the other 38."

Four differences separate them:

  • The second answer attaches a number and a unit to center and to spread. The first names the measures and stops.
  • The second says who was measured and what the variable is. The first could be about anything.
  • The second identifies the unusual features specifically: a gap in a stated interval, and 2 named high values. "A few outliers" locates nothing.
  • The second states the shape and then supports it by pointing at the long right tail.

Naming a measure is not reporting it

"The center is the median" names a statistic and stops. A reader still does not know whether a typical shopper spent 3 dollars or 300 dollars. The same problem sinks "the spread is the IQR" and "the shape is the shape of the histogram". Every feature needs a value attached to it.

Two weaker versions also fall short:

  • A value with no units. "The median is 15" leaves a grader choosing between dollars, minutes, and shoppers.
  • A value with no context. "The median is 15 dollars" is better, but "the median amount spent by these 40 shoppers is about 15 dollars" is the version that scores.

Numbers read off a histogram are approximate, and that is fine. Write "about 15 dollars", or say the median falls in the 10 to 20 dollar bin. A defensible estimate beats a precise-looking number you could not have known from the display.

Which center and spread pair to report

Pick the pair that matches the shape you just named.

  • Roughly symmetric with no outliers: report the mean xˉ\bar{x} (read "x-bar", the sum of the values divided by how many there are) and the sample standard deviation ss (a typical distance of a value from the mean).
  • Skewed, or outliers present: report the median and the interquartile range IQR=Q3Q1\text{IQR} = Q_3 - Q_1, where Q1Q_1 (the first quartile) and Q3Q_3 (the third quartile) bound the middle 50% of the data.

The reason is resistance. The AP course calls the median and IQR resistant, because outliers barely move them, and calls the mean, range, and standard deviation nonresistant. A single large value drags the mean upward and inflates the standard deviation while the median holds nearly still.

Reporting both pairs is allowed, but the pair you lean on should match the shape you just named. Leaning on the mean for a badly skewed distribution invites a follow-up about the outlier you did not mention, while the median and IQR describe the bulk of the data without that argument.

Unusual features: outliers, gaps, and clusters

The AP course defines these three precisely. An outlier is a value that is unusually small or large relative to the rest of the data. A gap is a region inside the distribution where no values were observed. A cluster is a concentration of values, usually set off from the others by a gap.

Modality belongs in the shape sentence. A distribution with one main peak is unimodal, one with two prominent peaks is bimodal, and one where every bar is about the same height with no prominent peak is approximately uniform. A bimodal shape often means two different groups got pooled into one graph, which is worth a sentence of its own.

When you have the raw data instead of a picture, you can back up an outlier claim with the 1.5 IQR rule rather than calling it by eye. From a histogram alone, "an apparent outlier" or "a value that stands apart from the rest" is the honest phrasing.

Where the points actually get lost

  • Skipping unusual features. This is the single most common omission, and it costs the same as skipping center.
  • Reversing the skew. The tail names the skew. A long right tail toward the large values is skewed right, no matter where the peak sits. Skewed left vs skewed right settles it.
  • Describing the graph instead of the data. "The bars go down as you move right" is not a shape claim. "The distribution is skewed to the right" is.
  • Dropping the units. Units are part of the context, and a number without them is incomplete.
  • Reporting the mean for a badly skewed distribution and then never mentioning that an outlier pulled it.
  • Answering a comparison question with one description. If two groups are shown, you owe a comparison, which is a different task. See how to compare two distributions.

Where this sits in the AP course

Describing distributions is topic 1.6 in Unit 1 of the Fall 2026 AP Statistics course, supported by topic 1.7 (summary statistics) and topic 1.8 (boxplots). Unit 1, Exploring One-Variable Data and Collecting Data, carries 20 to 30% of the multiple-choice section, and the skill being scored is 4.A, describing and comparing representations of data and summary statistics.

This shows up on free-response questions as often as on multiple choice, so practice writing the four sentences out. Work a set at describing distributions practice, or drag points and watch the shape change in the descriptive statistics sandbox. The official course description is at AP Central.

Describing a right-skewed distribution in context

A walk-in clinic records how many minutes each of 15 patients waited before seeing a nurse: 8, 22, 4, 12, 35, 7, 16, 5, 62, 11, 27, 8, 18, 10, 14. Write a complete description of the distribution.

  1. Sort the data: 4, 5, 7, 8, 8, 10, 11, 12, 14, 16, 18, 22, 27, 35, 62. There are n=15n = 15 values.

  2. Find the median. With 15 values the median is the 8th, so the median wait is 12 minutes.

  3. Find the quartiles using the median-excluded convention. The lower half is 4, 5, 7, 8, 8, 10, 11, whose middle value is Q1=8Q_1 = 8 minutes. The upper half is 14, 16, 18, 22, 27, 35, 62, whose middle value is Q3=22Q_3 = 22 minutes.

  4. Find the variability: IQR=Q3Q1=228=14\text{IQR} = Q_3 - Q_1 = 22 - 8 = 14 minutes, and the range is 624=5862 - 4 = 58 minutes.

  5. Check for outliers. 1.5×IQR=1.5×14=211.5 \times \text{IQR} = 1.5 \times 14 = 21, so the fences are 821=138 - 21 = -13 and 22+21=4322 + 21 = 43 minutes. The value 62 is above 43, so it is an outlier. Nothing falls below -13.

  6. Confirm the shape with the mean. The sum is 259 minutes, so xˉ=259/15=17.27\bar{x} = 259 / 15 = 17.27 minutes (rounded to 2 decimals). The mean sits well above the median of 12, which is what right skew does.

  7. Choose the summary pair. The distribution is skewed with a high outlier, so report the median and the IQR rather than the mean and standard deviation.

  8. Write the four sentences, each with a number, a unit, and the context.

The distribution of wait times for these 15 clinic patients is skewed to the right and unimodal, with a single peak among the shorter waits. The median wait is 12 minutes. The middle 50% of waits span an IQR of 14 minutes, from Q1=8Q_1 = 8 minutes to Q3=22Q_3 = 22 minutes, and the full range is 58 minutes. There is one high outlier at 62 minutes, which sits above the upper fence of 43 minutes and leaves a gap between it and the next longest wait of 35 minutes.

A symmetric distribution, where the mean and standard deviation belong

A cafe weighs 8 test doses from a new espresso grinder, in grams: 20, 18, 24, 19, 21, 16, 22, 20. Describe the distribution and justify which center and spread you report.

  1. Sort the data: 16, 18, 19, 20, 20, 21, 22, 24. There are n=8n = 8 values.

  2. Find the median. With 8 values it is the average of the 4th and 5th: (20+20)/2=20(20 + 20)/2 = 20 grams.

  3. Find the mean. The sum is 16+18+19+20+20+21+22+24=16016 + 18 + 19 + 20 + 20 + 21 + 22 + 24 = 160 grams, so xˉ=160/8=20\bar{x} = 160 / 8 = 20 grams. The mean equals the median, a signal of a symmetric shape.

  4. Find each deviation from the mean: -4, -2, -1, 0, 0, 1, 2, 4 grams. The deviations mirror each other, confirming symmetry.

  5. Square them and add: 16+4+1+0+0+1+4+16=4216 + 4 + 1 + 0 + 0 + 1 + 4 + 16 = 42.

  6. Divide by n1=7n - 1 = 7 to get the sample variance: 42/7=642 / 7 = 6 grams squared.

  7. Take the square root: s=6=2.45s = \sqrt{6} = 2.45 grams (rounded to 2 decimals).

  8. Check for outliers before committing to the mean. The lower half is 16, 18, 19, 20, giving Q1=(18+19)/2=18.5Q_1 = (18 + 19)/2 = 18.5; the upper half is 20, 21, 22, 24, giving Q3=(21+22)/2=21.5Q_3 = (21 + 22)/2 = 21.5. So IQR=3\text{IQR} = 3, 1.5×3=4.51.5 \times 3 = 4.5, and the fences are 18.54.5=1418.5 - 4.5 = 14 and 21.5+4.5=2621.5 + 4.5 = 26 grams. Every value falls inside, so there are no outliers.

The distribution of dose weights for these 8 test grinds is roughly symmetric and unimodal. The mean dose is 20 grams, matching the median of 20 grams. The standard deviation is about 2.45 grams, so a typical dose sits roughly 2.45 grams away from 20 grams, and the weights run from 16 to 24 grams (a range of 8 grams). There are no outliers, gaps, or clusters. Because the shape is symmetric with no outliers, the mean and standard deviation are the right pair to report here.

Frequently asked questions

Do I have to use SOCS?

No. SOCS (shape, outliers, center, spread) is a memory aid, not an answer. Writing the four letters or the four words with no values attached earns nothing. What gets scored is four statements, each with a number, a unit, and the variable named.

How precise do my numbers need to be if I only have a histogram?

Approximate is expected. Say the median falls in a stated bin, or give a value with "about" in front of it. Reporting a median of 14.7 dollars from a histogram with 10-dollar bins claims precision the display cannot support.

Do I need to mention outliers if there are none?

Yes. Write that there are no apparent outliers, gaps, or clusters. Silence about unusual features reads as an omission, and unusual features are the piece students skip most often.

Should I report the mean or the median?

Report the median and IQR when the distribution is skewed or has outliers, because both are resistant to extreme values. Report the mean xˉ\bar{x} and standard deviation ss when the distribution is roughly symmetric with no outliers. Reporting both is fine as long as the one you rely on fits the shape.

Is "bell-shaped" an acceptable shape description?

It is acceptable for a roughly symmetric, unimodal distribution, but say "roughly symmetric and unimodal" alongside it. Do not claim the distribution is normal from a small sample; normal is a model, and a histogram of 15 values cannot establish it.