What is a standardized test statistic?
By Jude Wallis · Published
A standardized test statistic counts how many standard errors your sample result sits from the value the null hypothesis claims: statistic minus null value, divided by the standard error of the statistic. Every z and t statistic in AP Statistics is that one formula with a different denominator.
AP Statistics: Unit 3 (topics 3.5 Setting Up a Test for a Population Proportion, 3.7 Carrying Out a Test for a Population Proportion). The same standardized form runs through both inference units of the Fall 2026 course: proportions in topics 3.5 to 3.7 and 3.12 to 3.13, means in topics 4.4 to 4.5 and 4.9 to 4.10. Chi-square for homogeneity and independence, topics 3.14 and 3.15, uses the squared-and-summed version instead.
The one formula behind every test statistic
Say the three pieces out loud, because each one has a job.
- The statistic is what your sample gave you: ("p-hat"), ("x-bar"), or a difference like .
- The null value is the number the null hypothesis claims for the matching parameter: , , or 0 when the null says two groups are the same.
- The standard error is the typical distance between the statistic and the parameter, from one random sample to the next.
What comes out is a count, in units of standard errors. A standardized test statistic of 2.3 says your sample landed 2.3 standard errors away from where the null hypothesis said it should land. That number, not the raw difference, is what becomes a p-value, because it is the version of the difference whose distribution is known.
Why standardizing is the whole idea
You met this move in Unit 1 as the z-score:
Subtract the center, divide by the spread, and you get a location that no longer depends on units. Metres, dollars and test points all come out as "how many standard deviations from the middle."
A test statistic performs the identical operation one level up, on the sampling distribution instead of on the data. The individual value becomes the statistic. The population mean becomes the value the null hypothesis claims for the parameter. And , the spread of individual values, becomes the standard error, the spread of the statistic across repeated samples.
That swap is the reason every procedure shares a shape. Raw differences are not comparable: a 0.03 gap in a proportion and a 12-point gap in a mean cannot be ranked against each other. Measured in standard errors, they can, and a single normal or curve turns either one into a probability. If you keep the difference between a standard deviation and a standard error straight, the rest of inference is bookkeeping.
What changes: the denominator
The numerator is always the same idea, so the work of learning inference is learning the denominator. Here is every one you need in this course.
| Procedure | Statistic | Null value | Standard error in the denominator |
|---|---|---|---|
| One proportion, | |||
| Two proportions, | 0 | ||
| One mean, | |||
| Paired mean, | 0 | ||
| Two means, | 0 |
One pattern explains the letter on the front. A proportion has a single parameter that controls both its center and its spread, so once the null names the standard error is a known number and the statistic is a . A mean has two independent parameters, and , and the null names only . You have to estimate with the sample standard deviation , and that estimate carries its own error, so the statistic is a . The full comparison is in t-test vs z-test.
What changes: the curve you compare it against
A standardized value is meaningless until you say which curve reads it. A goes to the standard normal curve. A goes to a curve with a stated number of degrees of freedom. The same number gives different answers.
| Curve | two-sided p-value for a statistic of 2.00 | |
|---|---|---|
| standard normal, | none | 0.0455 |
| 9 | 0.0766 | |
| 24 | 0.0569 | |
| 49 | 0.0511 | |
| 99 | 0.0482 | |
| 999 | 0.0458 |
A statistic of exactly 2.00 is significant at 0.05 read as a and not significant read as a with 9 degrees of freedom. The curves have heavier tails, which is the price of estimating from the data, and the tails thin out as grows until is indistinguishable from . That is why the row of the t-table sits within 0.005 of the row printed beneath it, 1.962 against 1.960 at 95% confidence.
So a complete answer names three things, not one: the statistic, the curve, and the degrees of freedom. "" is half an answer. " with " is the whole one.
Where chi-square fits, honestly
The chi-square statistic does not fit the template literally, and pretending otherwise causes real confusion. It is
Each term is the same idea, squared. is an observed count minus the count the null hypothesis predicts, which is a numerator of exactly the standard shape, and dividing by scales it, because for a count of that size is roughly its standard deviation. So is a sum of squared standardized gaps rather than one signed gap.
Two consequences follow from the squaring, and both are tested. Squaring throws away direction, so a chi-square test has no one-sided or two-sided version and its p-value is always the right tail. And because you sum over cells, the degrees of freedom count cells rather than sample size: for a two-way table.
The connection is exact in the smallest case. Run a two-proportion z-test on 45 successes out of 150 against 30 out of 150 and you get . Run a chi-square test for homogeneity on the same 2 by 2 table and you get , which is , with and the identical p-value of 0.0455. On a 2 by 2 table the two tests are the same test written twice. Details in chi-square tests explained.
The null value can appear in the denominator too
This is the detail that separates a test from an interval, and it is the most common place to lose an answer.
For a one-proportion test, the standard error is built from , not from . Everything in a test is computed as though the null hypothesis were true, and the null names the proportion, so it also names the spread: . For a one-proportion interval there is no null hypothesis in the room, so the only estimate available is , and the standard error is .
For a two-proportion test, the null says , so under the null there is one common proportion and both samples estimate it. Pool them:
and use in the standard error. For a two-proportion interval there is nothing to pool toward, so each sample supplies its own. On the 45-out-of-150 against 30-out-of-150 data, the pooled standard error is 0.050000 and the unpooled one is 0.049666: close, but they are answering different questions and the exam expects the right one in each place. See one-proportion vs two-proportion z-test.
Means have no equivalent split, because the null hypothesis about says nothing about . A statistic and a interval use the same .
Mistakes that cost points
- Dividing by the standard deviation instead of the standard error. describes the spread of the data; describes the spread of . Using makes the statistic times too small and can bury a real effect.
- Using in the standard error of a one-proportion test. Use . Save for the interval.
- Forgetting to pool for a two-proportion test. The null claims the proportions are equal, so the test uses one combined estimate.
- Reporting a statistic without its degrees of freedom. A with no cannot be turned into a p-value by the reader or by the scorer.
- Reading a statistic off the z-table. It gives a p-value that is too small, which can turn a non-significant result into a significant one.
- Confusing the statistic with the p-value. is a distance in standard errors. is a probability. They are never interchangeable.
One statistic of 2.00, two different curves
A random sample of 100 people gives 60 successes. Test against at . Then suppose a separate study produced a statistic of exactly 2.00 with . Compare the two p-values.
Compute the statistic: .
Build the standard error from the null value, because a test assumes is true: .
Standardize: . The sample proportion sits 2 standard errors above the claimed value.
Compare against the standard normal curve: . Since , reject .
Now take the same value, 2.00, to a curve with . The two-sided p-value is 0.0766.
Since , that study would fail to reject at the same significance level.
with a two-sided p-value of 0.0455, so you reject . The identical statistic read against a curve with gives 0.0766, so it would not be significant. The numerator and denominator are only part of the answer: the curve you compare against, and its degrees of freedom, decide the rest.
Two proportions: pooled for the test, unpooled for the interval
Group 1 has 45 successes in 150 trials, group 2 has 30 successes in 150. Test against at , then build a 95% confidence interval for . Finally run a chi-square test for homogeneity on the same 2 by 2 table and compare.
Statistics: and , so the observed difference is 0.10. The null value is 0.
The null claims the two proportions are equal, so pool: .
Pooled standard error: .
Standardize: , two-sided p-value , so reject .
For the interval there is no null value, so do not pool: .
Margin of error: , so the interval is , or 0.0027 to 0.1973.
Now the chi-square version. The table has row totals 150 and 150, column totals 75 and 225, grand total 300, so every expected count is or .
Add the four terms: .
With , the right-tail p-value for is 0.0455, and .
with , so reject . The 95% interval for is 0.0027 to 0.1973, built from a slightly different standard error (0.049666 rather than 0.050000) because an interval has no null value to pool toward. The chi-square test on the same table gives with and the same p-value, 0.0455, since on a 2 by 2 table.
Frequently asked questions
What is the difference between a test statistic and a standardized test statistic?
In practice they are used interchangeably for and . Strictly, a test statistic is any number computed from the data to measure evidence against the null hypothesis, and standardized names the particular form used here: statistic minus null value, over standard error. The chi-square statistic is a test statistic that does not take that literal form.
Why does the denominator use the standard error rather than the standard deviation?
Because the numerator is about a statistic, not about one observation. varies less from sample to sample than an individual value does, and is that smaller spread. Dividing by instead would compare the gap to the wrong ruler and make every test statistic about times too small.
Is the chi-square statistic a standardized test statistic?
Not in the literal form. It squares each observed-minus-expected gap, scales it by the expected count, and adds the terms up, so it measures a squared distance rather than a signed one. That is why its p-value is always the right tail. On a 2 by 2 table it equals the square of the two-proportion and gives the same p-value.
Why does a one-proportion test use in the standard error but an interval uses ?
A test computes everything as though the null hypothesis were true, and for a proportion the null value fixes the spread as well as the center, so belongs in the standard error. An interval starts from no claim at all, so the only estimate of the spread available is the one the data supply, .
Does every AP inference procedure use this formula?
Every and procedure does, for one proportion, two proportions, one mean, paired means and two means. Chi-square for homogeneity and independence, topics 3.14 and 3.15 in the Fall 2026 course, uses the squared-and-summed form instead.