Large Counts Condition vs 10% Condition
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Large Counts condition
Sampling distributions
The Large Counts condition checks that np is at least 10 and n(1-p) is at least 10, so the sampling distribution of p-hat is close to normal.
Large Counts is the shape check for proportion procedures: at least 10 expected successes and at least 10 expected failures before you may treat the sampling distribution of (p-hat, the sample proportion) as normal. It exists because is a binomial count divided by , and a binomial is symmetric only when , piling up against a wall as slides toward 0 or 1. Large Counts is the textbook name; the AP course calls the same check the normality condition.
Which value of goes in depends on the procedure. A significance test has a hypothesized value, so verify and using (p-naught). A confidence interval assumes nothing about , so verify and , which is just counting the successes and failures you observed. Testing with passes, since and . If those same 60 observations contained only 9 successes, the interval version fails on . One data set, two procedures, two different answers, and that is not a contradiction.
The misreading is " is more than 30, so the normal model is fine." The 30 belongs to sample means and never appears in a proportion problem. How large must be depends entirely on : at the condition demands , which is why the rule counts outcomes instead of counting observations.
Passing is not the same as exact. With and both counts clear the bar comfortably, yet the exact coverage of the nominal 95 percent z-interval is 94.1 percent. Break the condition at and , where , and coverage falls to 87.6 percent. The coverage simulator reports the exact coverage at any setting. The bar of 10 is a convention rather than a theorem, and some textbooks use 5 or 15, but AP standardized on 10.
10% condition
Sampling distributions
The 10% condition says a sample drawn without replacement should be no more than 10% of the population, so the observations stay nearly independent.
Sampling without replacement makes the draws dependent: taking one person out changes what is left for the next draw. The standard deviation formulas and are built for independent draws, and the exact spread under sampling without replacement carries an extra factor , where is the population size. The 10% condition, , is the promise that this factor is close enough to 1 to ignore.
Put right at the ceiling and the size of the approximation becomes visible. With and , the factor is , so the true standard deviation is about 5 percent below what the formula reports. The error runs in the safe direction, since a spread that is too large gives intervals slightly too wide rather than too narrow. Sample half the population, , and the factor falls to 0.7075: the true standard deviation is now 29 percent below the formula, which is no longer ignorable.
The sentence to correct is this one: "the sample has to be at least 10 percent of the population." It is backwards twice. The condition is a ceiling, not a floor, and it has nothing to do with being representative. A random sample of 1,000 from 300 million is about 0.0003 percent of its population and is perfectly sound, because representativeness comes from random selection rather than from the fraction sampled.
Keep the checks apart, because they do different jobs. Randomness comes from the design. Large Counts is about shape, whether a normal model fits the sampling distribution. This one is about spread, whether the independence formula is close enough to the truth. It is the check you skip when sampling is done with replacement, and the one you cite when arguing independence from a design.