Regression and correlation practice problems
By Jude Wallis · Published
These eight problems cover the core regression skills: finding the slope and intercept from summary statistics or raw data, interpreting the slope, r, and r-squared in context, and computing and reading residuals. Work each one first, then open the solution to check every step.
AP Statistics: Unit 5 (topics 5.2 Correlation, 5.3 Linear Regression Models, 5.4 Residuals, 5.5 Least-Squares Regression). This set spans Unit 5 of the Fall 2026 AP Statistics course: correlation (5.2), linear regression models (5.3), residuals (5.4), and least-squares regression (5.5). The redesigned course expects you to find coefficients with technology and interpret them in context, and it no longer includes inference for regression slopes.
What these problems build
These eight problems build the three regression skills Unit 5 tests together, using small clean numbers so you can focus on the method.
- Find the line. Compute the slope and intercept from the means and (the mean of the x-values and of the y-values), the standard deviations and , and the correlation , with and , or straight from paired data.
- Interpret in context. State the slope, the correlation , and the coefficient of determination in words, saying predicted and watching for extrapolation.
- Work with residuals. Compute a residual as observed minus predicted, (with read "y-hat"), and read what its sign says about the model.
For the method laid out step by step, read least-squares regression line, and check any line you build with the regression calculator. Try each problem before opening its solution, and keep extra decimals in , , and until the final step. For more sets like this one, see all practice problems.
Problem 1
A greenhouse trial links the weekly fertilizer dose (in grams) given to each tomato plant, , to the season's yield (in kilograms), . Across the plants the summary statistics are grams, kg, grams, kg, and correlation . Find the least-squares regression line and interpret the slope in context.
Show the worked solution
Recall the two formulas. The slope is , where and are the standard deviations (typical spread) of x and y and is the correlation. The intercept is , where ("x-bar") and ("y-bar") are the means of x and y.
Compute the slope. kg per gram.
Compute the intercept using the slope from the previous step. kg.
Write the line. .
Interpret the slope in context. The slope is kg per gram, so for each additional gram of weekly fertilizer, the model predicts the yield increases by 3 kg.
Check that the line passes through . At , , which matches .
The least-squares line is ; each additional gram of weekly fertilizer predicts a 3 kg increase in yield.
Problem 2
A university makerspace logs, for four one-hour blocks, the number of 3D printers running, , and the electricity used that hour in kilowatt-hours, : , , , . Find the least-squares regression line from the data.
Show the worked solution
Find the means. printers and kWh.
Find the deviations from each mean. For : , , , . For : , , , .
Sum the products of the paired deviations for the top of the slope. .
Sum the squared x-deviations for the bottom of the slope. .
Compute the slope. kWh per printer.
Compute the intercept. kWh.
Write the line. .
The least-squares line is , so each extra printer running predicts 1.6 more kilowatt-hours that hour.
Problem 3
A candle maker fit a least-squares line predicting a candle's burn time (in hours) from its weight (in ounces): , where is the weight and is the predicted burn time. (a) An 8-ounce candle actually burns for 15 hours. Find its residual and say what the sign means. (b) A 10-ounce candle has a residual of hours. Find its observed burn time.
Show the worked solution
Part (a): find the predicted burn time at . hours.
Compute the residual, observed minus predicted. hour.
Read the sign. The residual is positive, so the observed value sits above the line and the model underpredicted this candle's burn time by 1 hour.
Part (b): find the predicted burn time at . hours.
Rearrange the residual formula to observed equals predicted plus residual. hours.
Read the sign. The negative residual means the model overpredicted: this 10-ounce candle burned 2 hours less than the line predicted.
(a) The residual is hour, so the model underpredicted the burn time by 1 hour. (b) The observed burn time is hours, and the negative residual means the model overpredicted.
Problem 4
A cycling shop fit a least-squares line predicting a used e-bike's resale price (in dollars) from its odometer reading (in hundreds of kilometers), : . The e-bikes in the data had odometer readings from to (that is, 500 to 4000 km), and the correlation was . Interpret the slope, the intercept, the correlation , and the coefficient of determination in context.
Show the worked solution
Interpret the slope. Because is measured in hundreds of kilometers, the slope means that for each additional 100 km on the odometer, the model predicts the resale price drops by 50 dollars.
Interpret the intercept. The intercept is the predicted price at (0 km). Since the data start at 500 km, is outside the observed range, so this is an extrapolation, not a value read directly from the data.
Interpret the correlation. is close to , so there is a strong negative linear association: e-bikes with more kilometers tend to sell for less.
Compute the coefficient of determination. .
Interpret . About 81 percent of the variation in resale price is explained by the linear relationship with odometer reading.
Slope: each extra 100 km predicts a 50 dollar drop. Intercept: 1500 dollars predicted at 0 km, an extrapolation. is a strong negative association, and means about 81 percent of the price variation is explained by odometer reading.
Problem 5
A call center studies its shifts. Across many shifts the number of agents on duty, , has mean and standard deviation , while the calls resolved per hour, , has mean and standard deviation ; the correlation is . (a) Find the least-squares regression line. (b) One shift had 25 agents and resolved 80 calls in an hour. Find that shift's residual and interpret it.
Show the worked solution
Part (a): compute the slope. calls per agent.
Compute the intercept. calls.
Write the line. .
Part (b): find the predicted number of calls at . calls.
Compute the residual, observed minus predicted. calls.
Interpret the residual. It is positive, so the model underpredicted: this shift resolved 8 more calls than the line predicted for 25 agents.
(a) . (b) The residual is calls; with 25 agents the line predicted 72 calls, and the shift resolved 8 more than predicted.
Problem 6
An urban planner collects data on many neighborhoods and finds a correlation of between the number of coffee shops in a neighborhood, , and the average monthly apartment rent, . The data come from an observational study. (a) Interpret in context. (b) Compute and interpret the coefficient of determination . (c) Does this show that adding coffee shops causes rents to rise? Explain.
Show the worked solution
Part (a): interpret the correlation. is positive and moderately strong, so neighborhoods with more coffee shops tend to have higher average rent, and the association is roughly linear.
Part (b): compute . .
Interpret . About 49 percent of the variation in average rent is explained by the linear relationship with the number of coffee shops; the other 51 percent is not.
Part (c): weigh causation. The data are observational, not from an experiment, so a correlation on its own cannot establish cause.
Name a lurking variable. Something like neighborhood population density or overall desirability could raise both the number of coffee shops and the rent, producing the association without one causing the other.
(a) is a moderately strong positive linear association. (b) , so about 49 percent of the rent variation is explained by coffee-shop counts. (c) No; the study is observational, and a lurking variable such as neighborhood desirability could drive both.
Problem 7
A marine biologist measures sea turtles and summarizes shell length (in centimeters), , and body weight (in kilograms), : cm, cm, kg, kg, and . The measured shells ranged from 30 to 55 cm. (a) Find the least-squares regression line. (b) Interpret the slope. (c) Interpret the intercept, and explain any problem with it. (d) Find and interpret it. (e) Predict the weight of a turtle with a 50 cm shell, and state whether that prediction is interpolation or extrapolation.
Show the worked solution
Part (a): compute the slope, then the intercept. kg per cm, and kg. The line is .
Part (b): interpret the slope. The slope means that for each additional centimeter of shell length, the model predicts the weight increases by 1.6 kg.
Part (c): interpret the intercept. The intercept is the predicted weight at a shell length of 0 cm, which is a negative weight and therefore impossible. It also lies far outside the 30 to 55 cm range of the data, so it is an extrapolation with no logical meaning here.
Part (d): compute . , so about 64 percent of the variation in turtle weight is explained by the linear relationship with shell length.
Part (e): predict at . kg.
Classify the prediction. Because 50 cm lies between the smallest and largest measured shells, 30 and 55 cm, this is interpolation, so the prediction is reasonably reliable.
(a) . (b) Each extra cm of shell predicts 1.6 kg more weight. (c) The intercept kg is a negative weight and an extrapolation, so it has no real meaning. (d) , so 64 percent of the weight variation is explained. (e) kg, an interpolation since 50 cm is within 30 to 55 cm.
Problem 8
An ecologist runs a linear regression predicting a lake's dissolved oxygen (in mg/L), , from the water temperature (in degrees Celsius), . The output gives the line with correlation . (a) Interpret the slope. (b) At 20 degrees Celsius the measured dissolved oxygen was 8.5 mg/L; find and interpret the residual. (c) Find and interpret it. (d) The residual plot shows a clear U-shaped pattern. Given the high , what does that say about using a line here?
Show the worked solution
Part (a): interpret the slope. The slope means that for each additional degree Celsius, the model predicts the dissolved oxygen falls by 0.25 mg/L.
Part (b): find the predicted value at . mg/L.
Compute the residual, observed minus predicted. mg/L.
Interpret the residual. It is negative, so the observed value sits below the line and the model overpredicted the dissolved oxygen by 0.5 mg/L at 20 degrees.
Part (c): compute . , so about 90.3 percent of the variation in dissolved oxygen is explained by the linear relationship with temperature.
Part (d): read the residual plot. A U-shaped pattern is curvature, which is the direct sign that a line is not the most appropriate model, even though is high. A large does not confirm linearity; only apparent randomness in the residual plot does, so the true relationship here bends and the line misses in a systematic pattern.
(a) Each extra degree predicts a 0.25 mg/L drop. (b) Predicted 9 mg/L, so the residual is mg/L and the model overpredicted by 0.5. (c) , about 90.3 percent of the variation explained. (d) The U-shaped residual plot shows curvature, so a line is not appropriate despite the high .