Independence condition
By Jude Wallis · Updated
The independence condition requires that one observation gives no information about another, argued from random selection plus the 10% condition.
Independence means one observation carries no information about another, and it is a claim about the design rather than something you can read off the data. Sampling with replacement makes it exact. Sampling without replacement makes it false, because removing one person changes the odds for the next, so the course substitutes the 10% condition, , as an argument that the dependence is small enough to ignore. In a randomized experiment, random assignment supplies it instead.
The reason it is load bearing is one algebraic step. Variances of independent quantities add, so the sum has variance , and dividing that sum by divides its variance by , leaving as the variance of (x-bar, the sample mean). Take the square root and you have (sigma over the square root of n). With and , that is . The identical argument on a sum of independent successes produces for a proportion. Take independence away and the variance of the sum picks up covariance terms, the addition step fails, and the 2 is just a wrong number.
The sentence to unlearn is "the sample was random, so the observations are independent." Those are two different claims. A simple random sample drawn without replacement is random and dependent at the same time. Randomness is what centers the statistic on the parameter; the 10% condition is what makes the spread formula usable. Mislabeling one as the other is the error the course specifically flags.
Two boundaries. When data come from a randomized experiment the course treats the 10% check as unnecessary in the two-sample and chi-square procedures, because assignment rather than selection is doing the work. And every two-sample procedure needs a second independence claim stacked on this one, that the two groups are independent of each other, which is exactly what a matched pairs design gives up on purpose so that it can pair the observations instead.
More sampling distributions terms, or browse the full statistics glossary.