Residual vs Outlier
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Residual
Regression and correlation
A residual is the difference between an observed y-value and the value the regression line predicts for it: observed minus predicted.
A residual measures how far off the line is for a single point, so a positive residual sits above the line and a negative one below. It equals , the observed response minus the predicted response (read y-hat). For example, if the line predicts 72 but a student actually scored 80, the residual is . Least-squares regression chooses the line that makes the sum of the squared residuals as small as possible.
Outlier
Describing data
An outlier is a data value that lies unusually far from the rest of the distribution.
An outlier is a point that stands apart from the overall pattern of the data. A common rule flags a value as an outlier when it falls below or above , using the first quartile , the third quartile , and the interquartile range. For example, with , , and IQR , any value below or above is an outlier. Outliers can be genuine extreme cases or data-entry errors, so you investigate before removing them.