Describing data

The vocabulary for summarizing a single variable: center, spread, shape, and the graphs that show them.

37 terms

Bimodal distributionA bimodal distribution has two prominent peaks separated by a dip, marking two ranges where values cluster instead of one center.Center of a distributionThe center of a distribution is its typical value, reported with the mean when the shape is roughly symmetric and with the median when it is skewed.Cumulative relative frequencyCumulative relative frequency is the running proportion of the data that falls at or below a given value or ordered category.DeviationA deviation is the signed distance from a single data value to the mean, found by subtracting the mean from that value.DistributionA distribution describes which values a variable takes and how often each value or range of values occurs, across a data set or a population.Five-number summaryThe five-number summary reports five values in order: the minimum, first quartile, median, third quartile, and maximum of a data set.Interquartile range (IQR)The interquartile range (IQR) is the width of the middle half of the data, a single number equal to the third quartile minus the first quartile.IQR fence (1.5 IQR rule)An IQR fence is a boundary set 1.5 IQRs below the first quartile or above the third quartile, and values lying strictly beyond a fence count as outliers.Left-skewed distributionA left-skewed distribution has its long tail stretching toward the low values, which typically pulls the mean below the median.MeanThe mean is the arithmetic average of a set of values, found by adding them all up and dividing by how many there are. It is the balance point of the data.Mean absolute deviation (MAD)The mean absolute deviation is the average distance between the data values and the mean, using absolute values instead of squares.MedianThe median is the middle value of an ordered data set, splitting it so that half the values fall below and half above.MidrangeThe midrange is the average of the smallest and largest values in a data set, a quick measure of center that a single outlier can drag around.ModeThe mode is the value or category that appears most frequently in a data set.Modified boxplotA modified boxplot plots outliers as separate points and draws each whisker only to the most extreme value that is still inside the IQR fences.OutlierAn outlier is a data value sitting unusually far from the rest of the distribution, flagged by the 1.5 IQR rule or by a 2 standard deviation distance.PercentileA percentile is a value at or below which a given percentage of the data falls, so about 90 percent of values lie at or below the 90th percentile.Percentile rankA percentile rank is the percentage of values in a data set that fall at or below a given value, so it turns a raw score into a position within its group.QuartileA quartile is one of the three values that split an ordered data set into four groups of roughly equal size, marking the 25th, 50th, and 75th percentiles.RangeThe range is a measure of spread equal to the largest value in a data set minus the smallest, reported as a single number rather than as an interval.Relative frequencyA relative frequency is the count in a category divided by the total number of observations, giving that category as a share of the whole.Relative standingRelative standing is where a value falls within its own distribution, reported as a percentile, which is a rank, or a z-score, which is a distance.Resistant statisticA resistant statistic is a numerical summary whose value changes little when a few of the observations are extreme, so outliers cannot pull it far.Right-skewed distributionA right-skewed distribution has its long tail stretching toward the high values, which typically pulls the mean above the median.Shape of a distributionThe shape of a distribution is its symmetry or skew together with its number of peaks, read directly off a histogram, dotplot, or stemplot.SkewnessSkewness is the asymmetry of a distribution, named for the side its longer tail points toward rather than the side its tall bars sit on.Spread of a distributionThe spread of a distribution is how far apart its values are, summarized by the range, the interquartile range, or the standard deviation.Standard deviationThe standard deviation measures the typical distance of data values from the mean, and it is reported in the same units as the data itself.StandardizingStandardizing converts a value into a z-score by subtracting the mean and dividing by the standard deviation, putting different scales onto one common scale.Sum of squaresThe sum of squares is the total of the squared deviations from the mean, and it is the numerator that sits on top of the variance formula.Symmetric distributionA symmetric distribution has left and right halves that are approximate mirror images about its center, which puts the mean and the median together.Uniform distributionA uniform distribution spreads probability evenly, so every outcome or every interval of equal width is equally likely and the graph is flat.Unimodal distributionA unimodal distribution has a single clear peak, so its graph rises to one high point and falls away on both sides of it.VariabilityVariability is the tendency of values to differ, both among the observations in one data set and from one sample to the next.VarianceThe variance measures spread from squared distances to the mean, dividing their total by n - 1 for a sample and by n for a population.Weighted meanA weighted mean averages values after attaching a weight to each one, so values carrying larger weights pull the result further toward themselves.Z-scoreA z-score tells how many standard deviations a value lies above or below the mean of its distribution, so a negative z-score marks a value below the mean.