Correlation vs causation, with real examples

By Jude Wallis · Published

No. Correlation measures how two variables move together, but a strong link can come from a confounding variable, reverse causation, or coincidence. Only a randomized comparative experiment, which uses random assignment to balance out confounders, can establish that one variable causes another.

AP Statistics: Unit 5 (topics 5.2 Correlation, 1.13 Experimental Design, 1.10 The Investigative Question Revisited and Data Collection). In the Fall 2026 course, topic 5.2 states that correlation does not necessarily imply causation, while topic 1.13 explains that random assignment in an experiment is what allows a cause-and-effect conclusion.

Does correlation imply causation?

No, correlation does not imply causation. Correlation measures how strongly two quantitative variables move together in a straight-line pattern, reported as the correlation coefficient rr (a number between 1-1 and 11). A value of rr near 11 or 1-1 tells you the points fall close to a line, but it says nothing about why they line up.

The College Board states the rule directly: a real or perceived relationship between two variables does not mean that changes in one variable cause changes in the other. Put simply, correlation does not necessarily imply causation.

The strength of rr and the truth of a causal claim are separate questions. You can have a large rr with no causation at all, and a genuine causal effect can show a modest rr when the relationship is not perfectly linear.

Three situations produce a strong correlation without a direct causal link: a confounding (lurking) variable, reverse causation, or coincidence. You have to rule out all three before you claim that one variable drives the other.

Three reasons a correlation can mislead

A high rr is easy to find and easy to misread. These three mechanisms turn a real correlation into a false causal story, and confounding is the reasoning the design questions most commonly require.

Confounding (lurking) variables. In AP terms, a confounding variable is associated with both the explanatory variable and the response variable, so it supplies an alternative explanation for the pattern you see. Ice cream sales and drowning deaths climb together over the course of a year, yet neither causes the other. Warm weather is the confounder: heat drives ice cream sales up and also sends more people into the water, which raises the number of drownings.

Reverse causation. Sometimes the causal arrow points the opposite way from what you assumed. Towns with more police officers tend to report more crime, but hiring officers does not manufacture crime; heavy crime leads towns to hire more officers. Before locking in a direction, ask whether the response variable could be driving the explanatory variable instead.

Coincidence. Compare enough variables and some pairs will line up by chance with no mechanism connecting them. Over a single decade, the number of films an actor releases per year and the number of pool drownings per year can post a high rr that means nothing. A large rr from a short or cherry-picked series is weak evidence on its own.

What supports a causal claim: randomized experiments

To move from these two variables move together to this variable causes that one, you need a well-designed randomized comparative experiment. You impose the treatment yourself, include a comparison (control) group as a baseline, and use random assignment to sort experimental units into the groups.

Random assignment is the key step. It spreads confounding variables evenly across the groups on average, so the groups start out similar in every respect except the treatment. When that holds, a difference in the response can be attributed to the treatment rather than to some lurking variable. The College Board puts it this way: random assignment of treatments to experimental units allows cause-and-effect conclusions because the potential for confounding variables is reduced.

Observational studies, where you only record what already happened without assigning treatments, can reveal strong correlation but cannot rule out lurking variables. That is why an observational study supports association while a randomized experiment supports causation. See experiments vs observational studies for a side-by-side comparison.

How the AP exam tests correlation vs causation

The correlation vs causation question is a recurring AP theme. Correlation itself lives in Unit 5, topic 5.2, where you interpret rr and judge causal claims; the reliable move is to name a plausible confounding variable and state that the study design does not support causation.

The design half sits in Unit 1. Topic 1.13 (Experimental Design) explains that random assignment is what licenses a cause-and-effect conclusion, and topic 1.10 defines a confounding variable as one associated with both the explanatory and the response variable. Free-response questions often hand you an observational study and ask whether a causal conclusion is justified; unless treatments were randomly assigned, the answer is no.

Keep the two ideas paired: rr measures the strength of a linear relationship, not its cause. For how rr connects to the share of variation a model explains, see r vs r-squared. You can also read the official framework at AP Central.

Ice cream sales and drownings: a strong r that is not causal

One day is sampled from each of five months spread across a year. For each day you record ice cream sales xx (in hundreds of cones) and the number of drownings yy reported that day: (2,1)(2,1), (4,3)(4,3), (6,4)(6,4), (8,6)(8,6), (10,6)(10,6). Find the correlation coefficient rr and decide whether ice cream causes drownings.

  1. Find the sample means. Write xˉ\bar{x} for the mean of xx and yˉ\bar{y} for the mean of yy. xˉ=(2+4+6+8+10)/5=30/5=6\bar{x} = (2+4+6+8+10)/5 = 30/5 = 6 and yˉ=(1+3+4+6+6)/5=20/5=4\bar{y} = (1+3+4+6+6)/5 = 20/5 = 4.

  2. Find each deviation from the mean. For xx: 4,2,0,2,4-4, -2, 0, 2, 4. For yy: 3,1,0,2,2-3, -1, 0, 2, 2.

  3. Multiply the paired deviations and add them: (4)(3)+(2)(1)+(0)(0)+(2)(2)+(4)(2)=12+2+0+4+8=26(-4)(-3) + (-2)(-1) + (0)(0) + (2)(2) + (4)(2) = 12 + 2 + 0 + 4 + 8 = 26.

  4. Sum the squared xx deviations: 16+4+0+4+16=4016 + 4 + 0 + 4 + 16 = 40. Sum the squared yy deviations: 9+1+0+4+4=189 + 1 + 0 + 4 + 4 = 18.

  5. Apply the correlation formula $r=(xixˉ)(yiyˉ)(xixˉ)2(yiyˉ)2=264018=26720.r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2}\,\sqrt{\sum (y_i - \bar{y})^2}} = \frac{26}{\sqrt{40}\,\sqrt{18}} = \frac{26}{\sqrt{720}}.$

  6. Evaluate: 720=26.8328\sqrt{720} = 26.8328, so r=26/26.8328=0.9690r = 26 / 26.8328 = 0.9690.

  7. Interpret with the confounder. The correlation is strong and positive, but warm weather is a lurking variable that pushes both ice cream sales and swimming (and therefore drownings) up. The data are observational, so they cannot support the claim that ice cream causes drownings.

r0.969r \approx 0.969, a strong positive correlation, but it does not establish causation. Warm weather is a confounding variable associated with both quantities, and no treatment was randomly assigned.

Doctor visits and chronic conditions: reverse causation

For five adults you record annual doctor visits xx and number of chronic conditions yy: (2,1)(2,1), (3,1)(3,1), (4,2)(4,2), (5,3)(5,3), (6,3)(6,3). Find rr and explain why reading it as 'seeing the doctor causes chronic illness' is wrong.

  1. Find the means. xˉ=(2+3+4+5+6)/5=20/5=4\bar{x} = (2+3+4+5+6)/5 = 20/5 = 4 and yˉ=(1+1+2+3+3)/5=10/5=2\bar{y} = (1+1+2+3+3)/5 = 10/5 = 2.

  2. Deviations for xx: 2,1,0,1,2-2, -1, 0, 1, 2. Deviations for yy: 1,1,0,1,1-1, -1, 0, 1, 1.

  3. Paired deviation products, summed: (2)(1)+(1)(1)+(0)(0)+(1)(1)+(2)(1)=2+1+0+1+2=6(-2)(-1) + (-1)(-1) + (0)(0) + (1)(1) + (2)(1) = 2 + 1 + 0 + 1 + 2 = 6.

  4. Squared deviations: for xx, 4+1+0+1+4=104 + 1 + 0 + 1 + 4 = 10; for yy, 1+1+0+1+1=41 + 1 + 0 + 1 + 1 = 4.

  5. Apply the formula: $r=6104=640=66.3246=0.9487.r = \frac{6}{\sqrt{10}\,\sqrt{4}} = \frac{6}{\sqrt{40}} = \frac{6}{6.3246} = 0.9487.$

  6. Interpret the direction. The correlation is strong and positive, but the causal arrow runs backward: having more chronic conditions leads people to visit the doctor more often, not the reverse.

r0.949r \approx 0.949. The strong positive correlation reflects reverse causation, so more doctor visits do not cause chronic illness.

Frequently asked questions

Does correlation ever imply causation?

Not on its own. A correlation can be a clue that a causal link exists, but you can only conclude causation when a randomized experiment reduces the risk of confounding by balancing lurking variables across groups. Observational correlation, however strong, still leaves lurking variables and reverse causation on the table.

What is the difference between a lurking variable and a confounding variable?

A lurking variable is any variable outside your analysis that affects the results. It becomes a confounding variable when it is associated with both the explanatory and the response variable, which gives an alternative explanation for their relationship. On the AP exam, naming a plausible confounder is often enough to reject a causal claim.

Can a correlation coefficient near 1 prove causation?

No. An rr close to 11 or 1-1 only tells you the points fall near a straight line. It cannot tell you whether one variable drives the other, whether a third variable drives both, or whether the pattern is coincidence.