Is a higher standard deviation better? It depends

By Jude Wallis · Published

There is no universal answer. A higher standard deviation means more spread, and whether spread helps depends on what the numbers measure. Filling bottles or weighing a sample calls for a low standard deviation. A test meant to tell students apart is useless when the standard deviation is near zero.

AP Statistics: Unit 1 (topics 1.7 Summary Statistics for One Quantitative Variable, 1.9 Comparisons of the Distributions for One Quantitative Variable). Standard deviation is a summary statistic for one quantitative variable, Unit 1 topic 1.7 of the Fall 2026 AP Statistics course, and comparing spread between distributions is topic 1.9. Exam questions ask you to compare variability in context rather than to label a standard deviation good or bad.

The honest answer: it depends on what you are measuring

The standard deviation (ss for a sample, σ\sigma ("sigma") for a population) measures one thing: how far a typical value sits from the mean. It is a description of spread, not a score, and there is no threshold above which spread becomes bad or below which it becomes good.

The useful question is never "is this standard deviation high?" It is "what would go wrong if the values were more spread out, and what would go wrong if they were all identical?" Those two questions have different answers in different settings, and that is the entire reason this page cannot end with a rule.

Be suspicious of anyone offering one. A standard deviation of 5 is tiny for annual household incomes in dollars, enormous for the diameter of a machined bearing in millimeters, and meaningless without knowing which. Spread is judged against the unit, the mean, and the purpose.

When a lower standard deviation is better

Whenever the goal is to hit a target, consistency is the goal and spread is the enemy.

A filling machine. Two machines fill bottles labeled 500 mL. Both average exactly 500 mL, so the mean cannot tell them apart. Machine A has σ=2\sigma = 2 mL and Machine B has σ=8\sigma = 8 mL. If the acceptable range is 495 to 505 mL, then modeling each machine's output as approximately normal, Machine A puts about 98.76 percent of bottles in range and Machine B only about 46.80 percent. Same average, wildly different production line.

A measuring instrument. Weigh the same 100.0 g reference weight five times on two scales. A good scale returns 99.9, 100.0, 100.0, 100.1, 100.0, with a standard deviation of about 0.071 g. A worn scale returns 98.4, 101.3, 99.1, 101.7, 99.5, which also averages 100.0 g but has a standard deviation of about 1.432 g. Neither set of five readings gives evidence of bias, since both average 100.0 g, though five readings could not detect a small bias either way. Only one is precise enough to trust for a single reading.

Anything with a promise attached. Delivery times, dosages, the width of a runway, the time a surgery takes. In each case the average is a poor summary of the customer's experience, and reducing spread is the actual work.

When a higher standard deviation is better

The other direction is real, and it comes up more than students expect.

A test that has to separate people. Suppose eight students take Quiz A and score 80, 81, 81, 82, 82, 83, 83, 84. The mean is 82 and the standard deviation is about 1.31. The same eight take Quiz B and score 68, 72, 77, 81, 83, 87, 92, 96, again a mean of 82, with a standard deviation of about 9.59. Quiz B is the better instrument for ranking, because Quiz A compresses everyone into a five-point band from 80 to 84 where a single lucky guess reorders the class. Push it to the limit: if every student scored 82, the standard deviation would be 0 and the quiz would carry no information about who understood the material.

Choosing the conditions in a study. To learn how a fertilizer dose affects yield, you have to apply doses that differ. Testing 10 and 11 grams per plot gives you almost nothing to work with, while spanning 5 to 40 grams reveals the relationship. Deliberately building spread into the explanatory variable is standard design practice.

Anywhere the collection is supposed to span a range. A seed bank storing genetic diversity, a stock index tracking a whole market, a survey meant to capture the range of opinion in a town. When the collection is assembled deliberately, as with a seed bank, a small standard deviation is a sign the collector failed to cover the range. When the collection is meant to mirror something, as with an index or a survey, a small standard deviation is a finding about the thing being mirrored, not a fault in the collection.

The finance case, stated carefully

This example gets quoted often and usually gets mangled, so here is the accurate version.

In finance, the standard deviation of returns is the standard numerical measure of risk, often called volatility. Higher standard deviation means the returns bounce around more, which is not desirable in itself. Holding the average return fixed, a lower standard deviation is better. Combining investments that do not move in lockstep lowers the standard deviation of the combined return below the weighted average of the individual standard deviations, while the combined average return stays at the weighted average of the individual averages. That is the mathematical content behind the standard diversification argument.

The reason anyone tolerates a higher standard deviation is that historically, assets with higher average returns have also had higher volatility. So the choice is a tradeoff between two quantities, not a preference for spread. Saying "higher standard deviation is good for a portfolio" reverses the actual claim.

That is a description of how the statistic is used, not investment advice.

A standard deviation on its own tells you almost nothing

Before judging a standard deviation, you need three things it does not carry by itself.

  • The unit. A standard deviation of 5 is 5 of something. It only means what that something means.
  • The mean. For a variable measured on a scale with a true zero, comparing ss to xˉ\bar{x} ("x-bar") is a quick first look: a standard deviation of 5 next to a mean of 8 is large relative to the values, while the same 5 next to a mean of 5000 is small relative to them. The comparison stops working when the mean is near zero or when the zero point of the scale is arbitrary, and it never settles whether the spread is acceptable for the purpose.
  • A comparison. Spread becomes informative when it is compared to another group, to a specification, or to the same process last month. That is why AP questions ask you to compare distributions rather than to grade one.

One warning about a common mix-up. A larger sample does not shrink the standard deviation, because ss estimates how spread out the population is and that does not change when you look harder. What shrinks with a larger sample is the standard error, which measures how much a statistic like xˉ\bar{x} varies from sample to sample. Standard error vs standard deviation separates the two.

How to answer this on the AP exam

Exam questions rarely ask whether a standard deviation is good. They ask you to compare spread between distributions and to justify a choice in context, which is the answerable version of the question.

Three habits earn the points.

  • Say more or less variable, then say what that means here. "Machine B has a larger standard deviation, so its fill volumes are less consistent and more bottles fall outside the acceptable range" beats "Machine B has a higher standard deviation."
  • Use the resistant summary when outliers are present. If a distribution is strongly skewed or has an extreme value, compare the interquartile ranges instead, since the standard deviation gets dragged around by a single far-out value the way the mean does. Standard deviation vs IQR covers the choice, and why the median is not affected by outliers covers the mechanism.
  • Name the purpose. Whether more spread is a problem depends on what the process is for, so state that: filling bottles wants consistency, a placement test wants separation.

For the mechanics behind the number, see what standard deviation actually means, and for why the raw variance is not the thing you interpret, why variance is in squared units.

Two filling machines with the same mean

Two machines fill bottles labeled 500 mL. Machine A produces volumes that are approximately normal with mean 500 mL and standard deviation 2 mL. Machine B is approximately normal with mean 500 mL and standard deviation 8 mL. A bottle is acceptable if it holds between 495 and 505 mL. Which machine is better, and by how much?

  1. Notice that the means are identical, so the mean cannot distinguish the machines. The difference is entirely in the spread.

  2. Standardize the limits for Machine A. The z-score of 505 is z=5055002=2.5z = \frac{505 - 500}{2} = 2.5, and by symmetry the z-score of 495 is 2.5-2.5.

  3. Find the area between them: P(2.5<Z<2.5)=0.99380.0062=0.9876P(-2.5 < Z < 2.5) = 0.9938 - 0.0062 = 0.9876. About 98.76 percent of Machine A's bottles are acceptable, so about 1.24 percent are not, roughly 124 bottles in every 10,000.

  4. Standardize the limits for Machine B. The z-score of 505 is z=5055008=0.625z = \frac{505 - 500}{8} = 0.625, and the z-score of 495 is 0.625-0.625.

  5. Find that area: P(0.625<Z<0.625)0.73400.2660=0.4680P(-0.625 < Z < 0.625) \approx 0.7340 - 0.2660 = 0.4680. About 46.80 percent of Machine B's bottles are acceptable, so about 53.20 percent are not, roughly 5320 bottles in every 10,000.

  6. Compare in context. Machine B is on target on average and rejects more than half of what it makes.

Machine A is far better: about 98.76 percent of its bottles fall in the acceptable 495 to 505 mL range, against about 46.80 percent for Machine B. Here the lower standard deviation (2 mL vs 8 mL) is the whole difference, because the goal is hitting a target rather than covering a range.

Two quizzes with the same mean, and which one is more useful

Eight students take Quiz A and score 80, 81, 81, 82, 82, 83, 83, 84. The same eight take Quiz B and score 68, 72, 77, 81, 83, 87, 92, 96. Find the mean and sample standard deviation of each, then say which quiz is the better tool for placing students into an advanced section, and which is the better tool for certifying that everyone met a minimum standard.

  1. Quiz A sum: 80+81+81+82+82+83+83+84=65680 + 81 + 81 + 82 + 82 + 83 + 83 + 84 = 656, so xˉ=6568=82\bar{x} = \frac{656}{8} = 82.

  2. Quiz A squared deviations from 82: 4, 1, 1, 0, 0, 1, 1, 4, which add to 12. So s2=1271.7143s^2 = \frac{12}{7} \approx 1.7143 and s=1.71431.31s = \sqrt{1.7143} \approx 1.31.

  3. Quiz B sum: 68+72+77+81+83+87+92+96=65668 + 72 + 77 + 81 + 83 + 87 + 92 + 96 = 656, so xˉ=6568=82\bar{x} = \frac{656}{8} = 82 as well.

  4. Quiz B squared deviations from 82: 196, 100, 25, 1, 1, 25, 100, 196, which add to 644. So s2=6447=92s^2 = \frac{644}{7} = 92 and s=929.59s = \sqrt{92} \approx 9.59.

  5. For placement, the scores need to separate students. Quiz A spans only 80 to 84, so two students who differ by one question look nearly identical and small chance effects can flip the order. Quiz B spans 68 to 96, so the ranking is built on real gaps.

  6. For certifying a minimum, the goal flips. If the standard is 75, Quiz A shows every student clearing it comfortably, while Quiz B shows two students below 75, which may reflect the harder items rather than the students.

Both quizzes have a mean of 82. Quiz A has s1.31s \approx 1.31 and Quiz B has s9.59s \approx 9.59. Quiz B is the better placement instrument because its larger standard deviation separates students, while Quiz A is the more reassuring evidence that the whole class cleared a floor. The same statistic is a strength in one use and a weakness in the other.

Frequently asked questions

What counts as a good standard deviation?

There is no such number, and any rule of thumb you meet is smuggling in assumptions about the units and the purpose. A standard deviation only becomes interpretable next to the mean, the units, and something to compare it to, such as a specification limit or another group.

Is a high standard deviation bad?

Only when consistency is what you need. For a filling machine, a measuring instrument, or a delivery promise, a high standard deviation is a real problem. For a test that has to rank students or a sample meant to represent a varied population, spread is the point and too little of it is the failure.

What does a standard deviation of zero mean?

That every value in the data set is identical, so no value differs from the mean at all. Whether that is good depends entirely on context: perfect for a machine cutting parts to length, useless for an exam that is supposed to distinguish students.

Does collecting more data lower the standard deviation?

No. The standard deviation estimates how spread out the underlying population is, and that stays the same however much data you collect; more data just estimates it more reliably. The quantity that shrinks with a larger sample is the standard error of a statistic such as the sample mean.

Should I compare standard deviations or interquartile ranges?

Use standard deviations when both distributions are roughly symmetric with no extreme values, since it uses every observation. Use interquartile ranges when either distribution is strongly skewed or has outliers, because in a sample of reasonable size a single far-out value inflates the standard deviation while barely moving the IQR. In a very small sample even the quartiles can be dragged by the outlier.