Proportion vs Rate
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Proportion
Variables and data types
A proportion is a part-to-whole fraction between 0 and 1, found by dividing the count in a category by the total number of observations.
A proportion is a part of a whole. Write it , where (p-hat) is the sample proportion, is how many observations fall in the category and is how many observations there are altogether; the population version is written . Two properties follow straight from that definition and both are checks you can run. The numerator counts a subset of what the denominator counts, so . And the units cancel, so a proportion is a bare number with nothing attached.
In a neighborhood of 9,000 residents, 58 reported a bicycle stolen last year. The proportion is , which is 0.64 percent of residents.
"58 out of 9,000 is a theft rate of 6.44." The 6.44 is a real number and it is not the proportion. It is that proportion multiplied by 1,000, so the whole statement is 6.44 thefts per 1,000 residents per year, and dropping the base leaves a figure a thousand times its own meaning. The bound is the fastest check there is: a proportion of 6.44 would mean 644 percent of residents, so any part-to-whole answer above 1 is an arithmetic or labeling error rather than a finding. Convert on purpose. , and 0.006444, 0.6444 percent and 6.44 per 1,000 are three names for one quantity.
The part has to sit inside the whole. Change the wording from "58 residents reported a theft" to "58 thefts were reported" and stops being a proportion, because one resident can be robbed twice and the numerator is now counting events instead of people. That version is a rate, and unlike a proportion it has no ceiling of 1.
Unit 3, Inference for Categorical Data: Proportions, is built on as the estimator of , which are the sample and population versions of the same summary. See parameter vs statistic for that distinction on its own.
Rate
Variables and data types
A rate divides a count by the size of the base it came from, such as time or population, so it carries units like per year or per 1,000.
A rate is a count divided by the size of whatever produced it: a population, a stretch of time, a number of trips, an amount of exposure. Two things separate it from a proportion. Its numerator counts events, which need not be a subset of what the denominator counts, so a rate has no ceiling of 1. And it keeps its units, so a rate means nothing until you say per what. Every proportion can be restated as a rate, since 0.15 is 15 per 100 and 150 per 1,000. The reverse fails: no proportion equals 60 miles per hour.
The base decides the answer because it decides the question. Suppose country X records 1,200 road deaths in a year, with 6 million residents driving 90 billion vehicle miles, and country Y records 900, with 9 million residents driving 30 billion vehicle miles. Per 100,000 residents, X is at 20.0 and Y at 10.0. Per 100 million vehicle miles, X is at 1.33 and Y at 3.00. Both pairs are correct and they point opposite ways, because one answers how risky it is to live there and the other answers how risky a mile of driving is there.
"Country Y has the lower road death rate." There is no such thing as the rate, only a rate per a stated base over a stated window: Y is lower per resident and higher per mile driven. The window matters as much as the base, since 7 per 1,000 per year and 7 per 1,000 per decade differ by a factor of ten.
Plenty of quantities called rates are arithmetically proportions. An unemployment rate divides unemployed workers by the labor force; a response rate divides replies by people contacted. In both, the numerator sits inside the denominator and the value cannot pass 1. What makes 84 emergency room visits in a town of 12,000 a genuine rate, visits per person per year or 7 per 1,000 per year, is that one person can turn up twice.