Scatterplot vs Residual Plot
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Scatterplot
Graphs and displays
A scatterplot graphs paired values of two quantitative variables as points, showing the direction, form, and strength of their relationship.
A scatterplot shows one point per individual, with the explanatory variable on the horizontal axis and the response variable on the vertical. Both numbers must come from the same individual: two lists of equal length measured on different units do not make a scatterplot. You read it for direction, form, strength, and any point that departs from the pattern.
Five students study 1, 2, 3, 4 and 5 hours and score 62, 68, 74, 76 and 85. The cloud rises from left to right and looks close to linear. The least-squares line is (y-hat) and the correlation is .
The misreading is "the points look tight, so must be about 0.99." Strength cannot be read off reliably by eye, because how tight the cloud looks depends on the axis scaling and does not. Multiply every score by 3 and the plot stretches into a steeper, apparently tighter band, and is still 0.9859 to four places. Record hours in minutes instead and the plot flattens, and is unchanged again. Correlation is computed from standardized values, so multiplying either variable by a positive constant, or adding one to it, leaves the number alone while changing the picture completely.
A scatterplot also cannot report how many observations sit at one spot. Add four more students who each studied 3 hours and scored 74 and the plot looks identical, one dot at , while moves from 0.986 to 0.982. Overplotting like that can bury a whole subgroup in a dense region.
A correlation of zero is not the same as no relationship. The five points , , , and lie exactly on and give . The scatterplot shows that curve plainly, which is the reason to look at it before trusting any single summary number. Scatterplots are topic 5.1.
Residual plot
Regression and correlation
A residual plot graphs the residuals against the explanatory variable or predicted values, used to check whether a line fits the data well.
A residual plot puts on the vertical axis against either or on the horizontal axis, with a reference line at 0. Both horizontal choices are standard: plotting against rescales the horizontal axis, and reverses it when the slope is negative. You read the plot for shape, not for size: a bend says a straight line is the wrong model for the trend, and a fan says the scatter is not constant across the data.
Six students study 1, 2, 3, 4, 5 and 6 hours and score 60, 72, 65, 78, 71 and 86. Their least-squares line is , so the residuals plotted above through are -2, 6, -5, 4, -7 and 4. They land on both sides of zero, they do not bend, and they do not widen from left to right. With six points that is about as much as such a plot can support, so read it beside the scatterplot rather than on its own.
Here is the sentence to retire: "the residual plot has no upward trend, so hours and score are not associated." A residual plot from a least-squares fit can never show a linear trend. The correlation between the residuals and is exactly 0 by construction, for these six points and for every other data set, because the line already absorbed the straight-line part of the pattern. The plot answers whether the line has the right shape, not whether the variables are related.
A flat, patternless residual plot is also not proof that the model is correct. It only means nothing obvious is left over, and a small sample can hide a real curve. Curvature is still evidence that a linear model does not belong on the data, and the Fall 2026 course has no topic on transforming data to achieve linearity, so the expected response to a bend is to say a line is not appropriate rather than to re-express the variables.