Scatterplot practice problems with full solutions
By Jude Wallis · Updated
These 8 problems build one habit: describe a scatterplot completely, naming form, direction, strength, and unusual features in context. Along the way you choose explanatory and response variables, read predictions off a fitted line, and separate an outlier from a high-leverage point.
AP Statistics: Unit 5 (topics 5.1 Graphical Representations Between Two Quantitative Variables). Topic 5.1 opens Unit 5 of the Fall 2026 AP Statistics course, where you construct scatterplots from bivariate quantitative data, describe an association by its form, direction, strength, and unusual features, and justify a claim from the plot; Unit 5 is 10 to 20 percent of the multiple-choice section.
What these 8 problems build
Every problem here is about the picture itself, before any line is fitted or any correlation is computed. Two habits carry the whole set.
First, explanatory on the x-axis, response on the y-axis. The explanatory variable is the one whose values you use to explain or predict the other, and both variables have to be quantitative before a scatterplot is possible at all.
Second, describe all four things. A description of an association names the form (linear or non-linear), the direction (positive or negative), the strength (strong, moderate, or weak), and any unusual features such as clusters, gaps, or points away from the pattern. Answers that give two of the four are the most common way to lose credit on this topic. Say all four in context, and write that there are no unusual features when there are none rather than leaving the sentence out.
Because these problems are read rather than displayed, every one gives exact coordinates or the exact geometry of a drawn line, so sketch the points on graph paper before you start. Several problems hand you a straight guide line and ask how far each point sits from it, which is the same idea a residual formalizes once you fit a least-squares regression line. Problems 4 and 6 turn on the difference between a point with an odd value and a point with an extreme value, which you can watch happen in the influential point explorer.
Problem 7 has a strong association that no straight line should be used for, and problem 8 has a near-perfect straight line that proves nothing about cause, which is the trap correlation vs causation is built around. The matching course topic is 5.1 scatterplots, and the regression calculator will fit a line to any of these data sets once you have described the picture yourself.
Problem 1
A school data club is deciding what to plot. For each situation, say which variable is explanatory and which is the response, and which axis each belongs on. (a) A ferry operator records, for each of 30 crossings, the wind speed in kilometers per hour and the crossing time in minutes. (b) A dealer records, for each of 18 second-hand vans, the odometer reading in thousands of kilometers and the price the van sold for. (c) A coach records each runner's squad (junior or senior) and their 5 km finishing time. (d) A physiotherapist records each patient's forearm circumference in centimeters and their grip strength in kilograms, with no plan to predict either one from the other.
Show the worked solution
Start from the definition. The explanatory variable is the one whose values you use to explain or predict the other, and it goes on the x-axis; the variable being explained is the response and goes on the y-axis. A scatterplot also needs both variables to be quantitative and measured on the same individuals.
(a) Ask which direction the explaining runs. Wind slows a ferry down; a slow crossing does not change the weather. So wind speed is explanatory and goes on the x-axis, and crossing time is the response on the y-axis.
(b) Do the same for the vans. A van's mileage is fixed long before anyone sets its price, and buyers price a van partly by its odometer reading. Odometer reading is explanatory (x-axis) and selling price is the response (y-axis).
(c) Check the variable types before anything else. Squad is categorical: it records a label, junior or senior, not a number, so there is no coordinate to plot on an axis. You cannot make a scatterplot here; compare the two finishing-time distributions with side-by-side boxplots instead.
(d) Both variables are quantitative, so a scatterplot is possible, but neither one is naturally the explanatory variable. The roles come from the question you ask, not from the data. If the point is to predict grip strength from a tape measure, put forearm circumference on the x-axis and grip strength on the y-axis, and state that choice.
Note what the word explanatory does not claim. Calling a variable explanatory only says which one you are using to predict the other. It is not a claim that it causes the response, which is why problem 8 exists.
(a) Wind speed on the x-axis, crossing time on the y-axis. (b) Odometer reading on the x-axis, selling price on the y-axis. (c) Neither: squad is categorical, so no scatterplot is possible; compare the two time distributions instead. (d) Both are quantitative so a plot works, but there are no natural roles; pick the axis from the question, for example forearm circumference on the x-axis if you want to predict grip strength.
Problem 2
A repair shop tests 9 phones of one model and records the phone's age in months, , and the battery's measured capacity as a percent of its capacity when new, : , , , , , , , , . (a) Which variable is explanatory? (b) Sketch the points and describe the association completely. (c) A student answers only "strong and negative". What does that answer miss, and why does it matter?
Show the worked solution
(a) Age is fixed before the battery is tested and is what you would use to predict capacity, so age in months is explanatory and goes on the x-axis. Capacity is the response and goes on the y-axis.
(b) Build a guide line from the two end points so that strength is measurable rather than a guess. From to capacity falls points while age rises months, a slope of capacity point per month. That line is , and it checks out at both ends: and .
(b) Measure every point against the guide line.
Age Capacity Guide line Gap 3 97 97 0 6 94 94 0 9 92 91 12 88 88 0 15 86 85 18 81 82 21 80 79 24 75 76 27 73 73 0 (b) Read the form off the gaps. They never exceed 1 capacity point and they do not run positive at the ends and negative in the middle or the other way round, so there is no bend: the form is linear.
(b) Read the direction. As age increases, capacity decreases, so the direction is negative.
(b) Judge the strength against the size of the pattern. Capacity spans points across the plot while every point sits within 1 point of the guide line, so the scatter about the pattern is tiny compared with the overall change. That is a strong association.
(b) Check for unusual features. No point sits away from the pattern, there is no clustering, and the ages step evenly by 3 months with no gap, so there are no unusual features. Put together: there is a strong, negative, linear association between a phone's age and its battery capacity, with no outliers, clusters, or gaps.
(c) Compare that with the student's answer. It gives direction and strength but never names the form, never says whether anything unusual is present, and never mentions the two variables. A description is marked on all four features plus context, so an answer with two of them is an incomplete description of a plot that is otherwise easy to read.
(a) Age in months. (b) A strong, negative, linear association between a phone's age and its battery capacity: every point lies within 1 percentage point of the line , and there are no outliers, clusters, or gaps. (c) It names direction and strength but omits the form (linear) and the unusual features (none), and it never names the variables in context.
Problem 3
A county's winter contractor plots 12 storms from the past three seasons. Here is the snowfall in centimeters, ranging from 8 cm to 44 cm, and is the plow time in hours. The scatterplot is linear, positive, and strong, and the line drawn through the points passes exactly through and . (a) Find the equation of that line. (b) Predict the plow time after a 25 cm storm. (c) One 25 cm storm took 31 hours. How far from the line does it sit, and what does that position mean? (d) Interpret the slope in context. (e) The county asks for the predicted plow time after a 90 cm storm. Give the number the line produces and explain why you would not report it as a prediction.
Show the worked solution
(a) Find the slope from the two points the line passes through. hour per centimeter.
(a) Find the intercept from either point. Using : , so . The line is , where (read "y-hat") is the predicted plow time in hours. Check it against the other point: , which matches.
(b) Substitute . hours.
(c) Compare the actual time with the predicted time. hours, so this storm's point sits 5 hours above the line. The line underpredicted it: that storm took 5 hours longer to clear than its snowfall alone suggested, which points to something else about it, such as drifting or the time of day it fell.
(d) Interpret the slope with units attached. Each extra centimeter of snow predicts about 0.8 hour more plowing on average, and minutes, so a centimeter of snow buys the crew about 48 more minutes of work. Scaled up, 10 more centimeters predict about 8 more hours.
(e) Compute the number first. hours.
(e) Then locate 90 cm relative to the data. The largest storm on the plot is 44 cm, so 90 cm is more than double the range the line was built from. Nothing in the plot shows the pattern stays straight out there, and it probably does not: at that depth crews hit equipment limits and shift limits. The 78 hours is an extrapolation, so report it only with that warning, or decline to predict and say the data do not reach 90 cm.
(a) . (b) 26 hours. (c) hours above the line, so the line underpredicted this storm by 5 hours. (d) Each extra centimeter of snow predicts about 0.8 hour, or 48 minutes, more plowing on average. (e) The line gives 78 hours, but 90 cm sits far outside the observed 8 cm to 44 cm range, so it is an extrapolation and should not be reported as a prediction.
Problem 4
A bike-share operator plots its 10 docking stations, with the number of docks and the average number of rides started per day. Eight stations lie close to the line : , , , , , , , . The other two are station P at and station Q at . (a) Show that no station in the group of eight sits more than 2 rides from that line. (b) Which of P and Q is an outlier in the direction? (c) Which of the two has high leverage? Compute (read "x-bar"), the mean number of docks over all 10 stations, to support your answer. (d) Which of P and Q is more likely to change the slope of a fitted line, and why? (e) A new station R opens at . Classify it and say what it would do to the line.
Show the worked solution
(a) Predict at each of the eight x values from and subtract.
Docks Rides Line Gap 10 52 50 12 57 58 14 66 66 0 16 73 74 18 83 82 20 89 90 22 99 98 24 105 106 The largest gap is 2 rides, at the 10-dock station.
(b) Measure P and Q against the same line. For P, rides, but P has 130, so P sits rides above the line, which is 26 times the largest gap in the group of eight. For Q, rides, exactly what Q has, so Q sits on the line. P is the outlier in the direction; Q is not an outlier at all.
(c) Leverage is about the x value alone, so compute the mean number of docks. The eight ordinary stations give , then , so docks.
(c) Compare each point's distance from . Q sits docks above the mean, while the furthest of the others is the 10-dock station at below it. P sits at below the mean, close to the middle. Q is the high-leverage point and P has almost none.
(d) Separate the two ingredients. A point tilts the fitted line only when it is far from in x and also away from the pattern in y. P has a huge vertical miss but sits at the center of the x range, so it lifts the whole line slightly without tilting it much. Q has a powerful grip on the tilt, but it lands exactly on the pattern the other eight follow, so it holds the current slope in place instead of changing it. Neither one moves the slope much, which is the point: high leverage on its own is not influence.
(e) Classify R at by the same two tests. Its 42 docks put it even further from the mean than Q, so it has high leverage. The line predicts rides, but R has 60, a miss of rides below the line.
(e) Combine them. High leverage plus a large miss is exactly what makes a point influential, so least squares would swing the right end of the line down hard, flattening the slope and turning a tight positive picture into a scattered one. Before refitting, check the station: 42 docks with 60 rides a day suggests a station that just opened, was closed for works, or serves a location the rest of the plot does not represent.
(a) The largest gap is 2 rides, at the 10-dock station. (b) P: it sits rides above the line, while Q sits exactly on it. (c) Q: docks, and Q's 40 docks is 20.7 above that, further out than any other station. (d) Neither moves the slope much. P has a big miss but almost no leverage at 17 docks, and Q has huge leverage but no miss, so it anchors the existing slope. (e) R has high leverage and sits 118 rides below the line, so it is influential: it would drag the right end down and flatten the slope.
Problem 5
A bakery chain plots its 12 shops, with the floor area in square meters and the weekly transactions in hundreds. Six shops are on high streets: , , , , , . The other six are in retail parks: , , , , , . The plot the manager sees does not label which shop is which. (a) Describe the scatterplot completely. (b) Compute the mean point of each group of six and use the two mean points to state the overall drift of the plot. (c) Within each group, what is the direction of the association? (d) The manager says larger shops sell less, so the chain should stop opening large shops. Explain what is wrong with that reading. (e) What would you change about the plot?
Show the worked solution
(a) Start with the unusual features here, because they dominate the picture. The points fall into two clusters, one running from 30 to 60 square meters and one from 140 to 200 square meters, with a gap of 80 square meters in which no shop appears.
(a) Now the direction and form. The small-area cluster sits high, at 14 to 24 hundred transactions, and the large-area cluster sits low, at 9 to 15 hundred, so the overall drift across the plot is negative. The form is not one straight band: each cluster is its own narrow, roughly linear band, and the two bands are offset from each other, so a single straight line describes the plot poorly.
(a) Then the strength, stated honestly. Inside each cluster the points hug a line tightly, but across all 12 shops the association is only moderate, with about and about 0.40. Even that much strength comes from the empty gap between the two groups rather than from one relationship holding across the whole range of areas, so the plot looks tidier than the number is.
(b) Mean point of the high-street shops. , so square meters. , so hundred transactions.
(b) Mean point of the retail-park shops. , so square meters. , so hundred transactions.
(b) Compare the two mean points. Moving from to , area rises by square meters while mean transactions fall by hundred, about 717 transactions a week. That fall is the entire source of the plot's negative drift.
(c) Look inside each cluster. Among the high-street shops, transactions climb from 14 to 24 hundred as area goes from 30 to 60 square meters. Among the retail-park shops, they climb from 9 to 15 hundred as area goes from 140 to 200 square meters. Both clusters are positive, even though the plot as a whole drifts down.
(d) Name the error. The manager has read a difference between two kinds of location as though it were the effect of floor area. The plot contains two populations: high-street shops that are small and busy because of passing footfall, and retail-park shops that are large and quieter. Within either type, bigger goes with busier, so nothing here suggests that shrinking a retail-park shop would raise its transactions.
(e) Fix the plot before fixing the strategy. Record location type for every shop and plot the two groups with different symbols, or draw two separate scatterplots, then describe each association separately. Reporting one association across a plot with two obvious clusters hides the variable that is actually doing the work.
(a) Two clusters (30 to 60 and 140 to 200 square meters) separated by an 80 square meter gap; the overall drift is negative but only moderate, with about , and the form is two offset linear bands rather than one, so what strength there is comes mainly from the gap between the groups. (b) Mean points and : area rises 123.33 square meters while mean transactions fall 7.17 hundred, so the drift is negative. (c) Positive inside both clusters. (d) The manager is reading a difference between two location types as an effect of floor area; within each type, larger shops do more transactions. (e) Record and plot location type, with different symbols or two separate plots, and describe each group's association separately.
Problem 6
A tutoring center plots 7 students, with the number of sessions attended and the points gained between two practice tests: , , , , , , . Every point lies within 1 point of the line . (a) Describe the plot completely. (b) An eighth student attended 13 sessions and gained 4 points, having been ill on test day. How far from the line does that point sit, and where does it sit in the direction? (c) What does adding it do to the direction, to the apparent strength, and to a line fitted to all 8 points? (d) Suppose instead that same student had attended 7 sessions and gained 4 points. How would the effect on the fitted line differ, and why?
Show the worked solution
(a) Verify the claim about the line on a couple of points before leaning on it. At , against an observed 15, a gap of 1; at , against an observed 27, again a gap of 1; at , , exactly the observed value.
(a) Describe all four features. Form: linear, since the gaps stay within 1 point and show no bend. Direction: positive, since more sessions go with more points gained. Strength: strong, because gains rise from 6 to 36 across the plot while no point sits more than 1 point off the line. Unusual features: none, with sessions stepping evenly by 2 and no point away from the pattern.
(b) Find the vertical miss. At the line predicts points, and the student gained 4, so the point sits points below the line. That is roughly 30 times the largest gap among the original 7.
(b) Find its position in . The mean of the original sessions is sessions, so 13 sessions is 5 to the right of center and near the right edge of a range that runs from 2 to 14. The point has both a large vertical miss and real leverage.
(c) Direction: still positive, but visibly weaker. The new point sits very low at the right-hand end, exactly where the existing pattern is highest, so it argues against the upward trend rather than adding to it.
(c) Strength: clearly reduced. A plot in which every point was within 1 of a line now contains one point 29.5 below it, so the scatter about any single straight pattern is no longer small compared with the rise across the plot.
(c) Fitted line: least squares has to chase that point, and because the point is far to the right it pulls the right end of the line down. The slope flattens and the intercept rises, so this one student is an influential point, not just an odd result.
(d) Redo both measurements at . The line predicts points, so the miss is points below the line, still far outside the pattern. But 7 sessions sits at the center of the x values, one below , and a point at the center has almost no leverage.
(d) Read off the difference. From the middle the point presses the whole line down instead of tilting it: the intercept falls, the slope barely moves, and the plot still looks positive. The same unusual value does very different damage depending on its value, and distance from is what gives a point the grip to change the slope.
(a) A strong, positive, linear association between sessions attended and points gained, with every point within 1 of and no outliers, clusters, or gaps. (b) It sits points below the line, at 13 sessions, which is 5 above and near the right edge of the data. (c) Still positive but much weaker, and the fitted line's right end is dragged down, so the slope flattens and the intercept rises: the point is influential. (d) At 7 sessions the miss is points, but the point sits at the center of the x values, so it pushes the whole line down without tilting it and the slope barely changes.
Problem 7
A cold-brew producer steeps the same coffee for different lengths of time and measures the caffeine in the finished brew. With the steep time in hours and the caffeine in milligrams per 100 mL, the readings are , , , , , , , . (a) Find the change in caffeine over each 2-hour step and describe what those changes show. (b) Describe the scatterplot completely. (c) The producer's analyst draws the straight line through the first and last points. What does that line predict at 6 hours and at 8 hours, and how far off is it? (d) If a least-squares line were fitted anyway, what pattern would the residuals show? (e) The analyst uses the straight line to predict the caffeine after 24 hours. Compute that prediction and explain why it cannot be right.
Show the worked solution
(a) Subtract each reading from the next: , , , , , , mg per 100 mL.
(a) Read the sequence of changes. Every step is positive, so caffeine keeps rising, but each step is smaller than the one before it. As a rate, the first step gains mg per 100 mL per hour and the last gains mg per 100 mL per hour, so the rate of increase falls by a factor of 14 across the plot. A straight line claims one constant rate, and no single rate describes this.
(b) Describe all four features. Form: non-linear, a smooth curve that rises steeply at first and flattens toward the right, so it is concave down. Direction: positive throughout, since caffeine never falls as steep time increases. Strength: strong, because the points follow that one curve with almost no scatter. Unusual features: none, with no clusters, no gaps, and no point off the curve.
(b) Note the trap in the wording. Strong does not mean linear. This association is about as strong as an association gets and is still the wrong shape for a line, so a description that says "strong positive linear" would be wrong on the first word that matters.
(c) Find the chord through the end points. Its slope is mg per 100 mL per hour, so the line is .
(c) Evaluate it at 6 hours. mg, against an observed 88 mg, so the line sits mg too low.
(c) Evaluate it at 8 hours. mg, against an observed 102 mg, so the line sits mg too low. Both middle values tower over the chord, which is what a concave-down curve does: the chord passes under every point between its ends.
(d) Predict the residual pattern. A fitted line would run above the data at both ends and below it through the middle, so the residuals would be negative at the ends and positive in the middle, forming a clear arch. That is a systematic pattern rather than random scatter, and it is the standard signal that the form of the model is wrong rather than the data being noisy.
(e) Compute the 24-hour prediction. mg per 100 mL.
(e) Compare it with what the plot is doing. Caffeine has already flattened near 123 mg by 16 hours, gaining only 2 mg over the last 2-hour step, and the grounds hold a finite amount of caffeine to extract. Reaching 170 mg would need the early steep rate to return and hold for another 8 hours, which the data contradict. The prediction commits two errors at once: a straight line for curved data, and extrapolation past the longest steep time observed.
(a) Steps of mg per 100 mL: still rising, but the rate falls from 14 mg per 100 mL per hour to 1 mg per 100 mL per hour. (b) A strong, positive, non-linear (concave down) association between steep time and caffeine, with no clusters, gaps, or points off the curve. (c) The chord has slope mg per 100 mL per hour and predicts 63.71 mg at 6 hours (24.29 mg low) and 75.57 mg at 8 hours (26.43 mg low). (d) An arch: residuals negative at both ends and positive through the middle, the signature of a wrong form. (e) It gives 170.43 mg, which is impossible here: the curve has flattened near 123 mg and 24 hours is beyond the data.
Problem 8
A city council plots its 8 policing neighborhoods, with the number of bus stops and the number of reported bicycle thefts in a year: , , , , , , , . Every point lies within 1 theft of the line . (a) Name the explanatory variable and describe the association completely. (b) Use the line to predict the thefts in a neighborhood with 40 bus stops, and say whether that prediction is safe. (c) A councillor proposes removing bus stops in order to cut bicycle theft. Explain why this plot cannot support that. (d) The council also records residents: the 25-stop neighborhood has 9,000 and the 50-stop neighborhood has 19,000. Compute thefts per 1,000 residents for each and say what that does to the councillor's argument. (e) What would it take to establish that bus stops cause thefts?
Show the worked solution
(a) The council is using bus stops to explain thefts, so the number of bus stops is the explanatory variable on the x-axis and reported thefts is the response on the y-axis. Check the line on two points: at , , exactly the observed value; at , against an observed 123, a gap of 1.
(a) Give all four features in context. Form: linear, since the gaps stay within 1 theft with no bend. Direction: positive, since neighborhoods with more bus stops report more thefts. Strength: strong, because thefts run from 33 to 123 across the plot while no point is more than 1 theft off the line. Unusual features: none, with no clusters, no gaps, and no point away from the pattern.
(b) Substitute . thefts. The observed bus-stop counts run from 12 to 58, so 40 is inside that range: this is an interpolation and is reasonable to report as a prediction.
(c) Separate what the plot shows from what the councillor wants. These are 8 neighborhoods observed as they are, with nobody assigning bus stops to them, so this is observational data. A tight straight-line pattern describes how two counts move together; it says nothing about what would happen if you intervened and pulled stops out.
(d) Convert both counts to rates. For the 25-stop neighborhood, , which is thefts per 1,000 residents. For the 50-stop neighborhood, , which is thefts per 1,000 residents.
(d) Read what happened. The neighborhood with twice as many bus stops has the lower theft rate, against per 1,000 residents. The strong positive pattern in the counts was, at least in part, a picture of how large each neighborhood is: more residents means more bus stops and also more bicycles being ridden and parked. Population is a lurking variable driving both counts, and once you account for it the councillor's association weakens rather than holds.
(e) Say what would settle the question. Only an experiment can, in principle: randomly assign which neighborhoods lose bus stops, hold other policing and infrastructure changes fixed, then compare theft counts afterwards. That is expensive and would be hard to justify to residents who depend on the buses.
(e) Give the realistic alternative. Compare rates rather than raw counts, and compare neighborhoods matched on population, density, and bicycle ownership so that size is not doing the explaining. A near-perfect straight line in an observational plot is still not evidence of cause, no matter how strong it looks.
(a) Bus stops is explanatory. There is a strong, positive, linear association between the number of bus stops and reported bicycle thefts, with every point within 1 theft of and no outliers, clusters, or gaps. (b) thefts, and 40 is inside the observed 12 to 58 range, so the prediction is an interpolation and safe to report. (c) The data are observational and the plot shows association only, so it cannot say what removing stops would do. (d) and thefts per 1,000 residents, so the neighborhood with more stops has the lower rate; population is a lurking variable inflating both counts. (e) A randomized experiment assigning stop removals, or failing that, comparing rates across neighborhoods matched on population, density, and bicycle ownership.