Residual vs Outlier

Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.

Residual

Regression and correlation

A residual is the difference between an observed y-value and the value the regression line predicts for it: observed minus predicted.

A residual measures how far off the line is for a single point, so a positive residual sits above the line and a negative one below. It equals residual=yy^\text{residual} = y - \hat{y}, the observed response minus the predicted response y^\hat{y} (read y-hat). For example, if the line predicts 72 but a student actually scored 80, the residual is 8072=880 - 72 = 8. Least-squares regression chooses the line that makes the sum of the squared residuals as small as possible.

Full entry for residual

Outlier

Describing data

An outlier is a data value that lies unusually far from the rest of the distribution.

An outlier is a point that stands apart from the overall pattern of the data. A common rule flags a value as an outlier when it falls below Q11.5×IQRQ_1 - 1.5 \times \text{IQR} or above Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR}, using the first quartile Q1Q_1, the third quartile Q3Q_3, and the interquartile range. For example, with Q1=3Q_1 = 3, Q3=9Q_3 = 9, and IQR =6= 6, any value below 39=63 - 9 = -6 or above 9+9=189 + 9 = 18 is an outlier. Outliers can be genuine extreme cases or data-entry errors, so you investigate before removing them.

Full entry for outlier

Where each one fits in the course