Symmetric Distribution vs Normal Distribution
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Symmetric distribution
Describing data
A symmetric distribution has left and right halves that are approximate mirror images about its center, which puts the mean and the median together.
Symmetry is a mirror test. Fold the graph at its center and a symmetric distribution has its two halves land on each other, which means a value a given distance above the center is about as common as one the same distance below. For real data the honest word is approximately symmetric; exact symmetry belongs to models such as the normal curve. When a distribution is symmetric, the mean ("x-bar") and the median sit in the same place.
The six values 3, 5, 7, 7, 9, 11 are symmetric about 7: the distances from 7 are , , , , , , each matched by its opposite. The mean is and the median is . Now take 1, 1, 2, 9, 10, 10. That set is just as perfectly symmetric, about 5.5, with two clumps and a hole where its center is.
The second set kills the sentence "it looks symmetric, so use the empirical rule." Symmetric is not the same as normal. For 1, 1, 2, 9, 10, 10 the mean is 5.5 and the sample standard deviation is about 4.59, so one standard deviation either side spans roughly 0.91 to 10.09 and captures all six values, 100 percent rather than the 68 percent the empirical rule would predict. That rule needs the bell shape, and symmetry alone does not supply it.
The implication runs one way only. Symmetric gets you the mean equal to the median; equal centers do not get you symmetry. For 1, 2, 4, 4, 9 the mean and the median are both 4, and yet the largest value sits 5 above that center while the smallest sits only 3 below. Read symmetry off a graph, then check the centers, not the reverse.
Judge symmetry from a display with equal-width bins, and remember that a dozen values wobble too much for the word to be more than a description. Ask whether the departure from a mirror image is bigger than the bumpiness that sample size would produce anyway. Shape vocabulary is Unit 1 topic 1.6.
Normal distribution
Random variables and distributions
The normal distribution is a continuous bell-shaped model in which probability is area under a curve fixed entirely by the mean and the standard deviation.
A normal distribution is a continuous model written : a density curve, symmetric about its mean (mu), single-peaked, and spread out by its standard deviation (sigma). Two numbers fix the whole curve. Probability is area underneath it, the total area is exactly 1, and half of that area sits on each side of . Because the area over a single point is zero, every normal question is really a question about an interval.
Suppose adult male heights are approximately in inches. To find the share above 72 inches, standardize: . The area to the left of is 0.8577, so about 0.142, roughly 14 percent, are taller than 72 inches. The same curve puts about 71.6 percent of men between 66 and 72 inches.
The claim that ruins the most work is "the sample is large, so the data are normal." Sample size does not change the shape of the variable being measured. A right-skewed variable such as household income stays right-skewed however many households you collect. What a large sample buys is that the sampling distribution of (x-bar, the sample mean) is close to normal, and that is a statement about the average of a sample, not about the individual values inside it.
A normal curve also runs on forever in both directions, so a normal model always assigns some probability to values the real variable cannot reach, negative heights included. For heights that leftover is far too small to matter. Where the model genuinely fails is shape: strongly skewed, hard-bounded, or clearly bimodal data should not be pushed through a normal calculation, and no amount of extra data repairs that.
The normal distribution is topic 2.11 in Unit 2, Probability, Random Variables, and Probability Distributions. The empirical rule gives the quick version of its areas and the z-table gives the exact ones.
There is a calculus reason the areas have to come from a table or a calculator rather than from an antiderivative. The bell-shaped function has no elementary antiderivative at all, so no amount of algebra produces a formula for the area under it between two bounds. CalcLearn works through that function's behaviour in the limit of e to the minus x squared.