Observational Study vs Census

Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.

Observational study

Collecting data and study design

An observational study measures individuals without assigning treatments, so it can reveal associations but cannot by itself establish cause and effect.

In an observational study the researcher records variables as they already stand and imposes nothing. Nobody is assigned to a group. The groups are formed by what the subjects already do or already are, and that single fact is why the design cannot settle cause. A survey is an observational study run by asking a standard set of questions.

An employer finds that 200 of its 500 staff use a standing desk and 300 do not. The standing desk users report a mean of 3.1 days of back pain per month against 4.6 days for everyone else, a difference of 1.5 days. That difference is real and it is an association. It is not evidence that the desks did anything, because the staff sorted themselves, and whoever chose a standing desk may also be younger, more active, or in better health to begin with. Every one of those traits travels with the group.

So this sentence fails: "the study shows that standing desks reduce back pain." The version that survives is "staff who use standing desks reported fewer days of back pain than staff who do not." The verbs carry the whole difference. Notice too that the arrow could point the other way. If the staff with chronic pain were the ones who asked for standing desks, this same design would show an association for a reason that has nothing to do with the desks working.

The reason is specific enough to write down. A confounding variable is one whose effect on the response cannot be separated from the explanatory variable's, and an observational study contains no mechanism that balances them. Random assignment balances the extraneous variables nobody thought to measure, which is its entire purpose; statistical adjustment can only handle the ones that were measured and recorded. A larger observational study estimates the association more precisely and does not make it one bit more causal.

Random selection is still worth having here. It earns the right to generalize the association to the population sampled, which is a separate permission from causation and is worked out under scope of inference.

Full entry for observational study

Census

Collecting data and study design

A census collects data from every individual in a population, so the value it produces is the parameter itself rather than an estimate of it.

In a census you record data on every individual in the population, so what you compute is the parameter and not an estimate of it. There is no sampling distribution around it, no standard error, no margin of error, because nothing was sampled. What makes a study a census is coverage of the population as you defined it, which means the same set of responses can be a census of one group and a sample of a larger one.

A department has 8 teachers with 3, 5, 6, 8, 11, 12, 14, and 21 years of experience. If the population is that department, the mean 808=10\frac{80}{8} = 10 years is μ\mu (mu), exact and final: no interval is needed because no other value is possible. Ask instead about all teachers in the district and those 8 become a sample, the same 10 years becomes xˉ\bar{x} (x-bar), and it turns into an estimate carrying uncertainty. The arithmetic did not move. The population did.

"A census has no error" holds for exactly one kind of error. It removes sampling error, the sample-to-sample variation that exists because you measured a part instead of the whole. Every other flaw survives: households an enumerator never reaches are undercoverage, people who refuse at the door are nonresponse, and a leading question produces the same response bias at full coverage that it produces in a sample of 300.

There is also nothing left to infer. When the data cover the population of interest, a confidence interval has no unknown parameter to bracket and a significance test has no population claim to weigh. If two fully measured departments average 10 and 11.4 years of experience, that 1.4 year gap is a description of those two departments, and testing it for significance answers a question nobody needed to ask.

A census is rare for practical reasons rather than statistical ones: cost, time, populations that change while you count them, and measurements that destroy what they measure, since testing every battery until it fails leaves none to sell.

Full entry for census

Where each one fits in the course