Power vs Type II error (power = 1 - beta)
By Jude Wallis · Published
A Type II error is failing to reject a false null hypothesis, with probability beta. Power is rejecting it correctly, with probability 1 minus beta. They are one fact stated from opposite ends. Neither is a single number for a test: both are computed against one specific true parameter value.
AP Statistics: Unit 3 (topics 3.8 Potential Errors When Performing Tests). Power and the Type II error rate belong to Unit 3 topic 3.8, Potential Errors When Performing Tests, in the Fall 2026 AP Statistics course. The z test used throughout has a known sigma, so every figure on this page can be checked against a z table by hand, up to the rounding of the critical value.
Power vs Type II error: the short answer
Both describe what happens when the null hypothesis is false. A Type II error is failing to reject it anyway, and its probability is ("beta"). Power is the other outcome, correctly rejecting a false , with probability . The two probabilities add to 1 because the two outcomes are the only two available. There is one quantity here, not two.
So the interesting part is not the difference. It is the qualifier both of them carry, and it is the one students drop. Neither nor power is a property of a test on its own. Each is computed by naming one specific value the parameter might really have, and each changes the instant you name a different one. The same test has one power against and a different power against . A power quoted with no alternative value attached is not a number anyone can check.
Neither one is a single number for a test
Take the z test the site's power visualizer runs, chosen because the population standard deviation is known, so nothing is estimated on the way to the answer and every number below can be checked against a z table.
Test against , with ("sigma") known to be 15 and . The standard error is , and at the test rejects once the sample mean ("x-bar") clears . That is the four-decimal critical value; a z table prints it as 1.645, which shifts the cutoff to 104.935 and every probability below by less than 0.0001.
Nothing in that setup mentions the truth. Supply one and the two numbers appear.
| If the true is | Power | Chance of failing to reject |
|---|---|---|
| 100 | 0.0500 | 0.9500 |
| 102 | 0.1640 | 0.8360 |
| 105 | 0.5087 | 0.4913 |
| 108 | 0.8466 | 0.1534 |
| 110 | 0.9543 | 0.0457 |
The test never changed. The cutoff sat at 104.9347 in every row, stayed at 0.05, and stayed at 25. Only the assumed truth moved, and power went from the significance level itself to near certainty. Read the last two columns and you have read power and at the same time, which is the whole relationship.
One caution about the first row. The right-hand column is in every row except that one, where makes true. A Type II error cannot happen when the null is correct, so 0.9500 there is the chance of a correct decision, not an error rate.
The two boundaries that catch people out
Two rows of that table deserve their own note, because both contradict a rule students invent.
Power equals when the truth sits at the null value. At power reads 0.0500, exactly the significance level. With no effect to detect, the only rejections still available are Type I errors, so the rejection rate is by construction. Change to 10 or to 200 and that row does not move.
Power can fall below . This test looks upward only, so it is blind to a true mean under 100. Against its power is 0.0005, roughly a hundredth of . The two-sided version of the same test cannot do that: with split across both tails the power curve bottoms out at 0.0500 when and climbs in both directions, reaching 0.3848 against . The two-sided test pays for that coverage elsewhere, since against its power is 0.3848 rather than 0.5087. That is one concrete price of the one-sided or two-sided decision.
The four levers, and which way each one pushes
Four things move power. Every direction below was recomputed one lever at a time, holding the other three fixed, from the same baseline: the test above aimed at a true mean of 105, where power is 0.5087 and is 0.4913.
| Lever | Change | Power | Beta | What it costs |
|---|---|---|---|---|
| Sample size | 25 to 100 | 0.5087 to 0.9543 | 0.4913 to 0.0457 | More data, and stays at 0.0500 |
| Significance level | 0.05 to 0.10 | 0.5087 to 0.6499 | 0.4913 to 0.3501 | The Type I error rate doubles |
| Distance of the truth from the null | 105 to 108 | 0.5087 to 0.8466 | 0.4913 to 0.1534 | Nothing, and it is not yours to set |
| Variability | 15 to 10 | 0.5087 to 0.8038 | 0.4913 to 0.1962 | A cleaner measurement process or a design that removes noise, and stays at 0.0500 |
The cost column is where they differ, and it is what a good exam answer names. Two of them buy power outright, and they buy it the same way: raising and lowering both shrink the standard error while leaving exactly where you set it. Raising is a trade, not a gain: it lowers and raises the false-alarm rate in the same move, which is the judgment covered in how to choose a significance level. The distance from the null is not a lever at all, only a fact you discover afterward. The route is the one students forget on free-response: a tighter measurement process, or a design that strips out a source of variation, raises power with no extra data and no extra false alarms; see standard error vs standard deviation for why those two are not the same quantity.
Say which lever you mean. "Increasing the sample size increases power" is true only with the other three held fixed, and dropping that qualifier is the most common way a correct sentence turns into a wrong one.
Power vs Type II error side by side
| Feature | Power | Type II error |
|---|---|---|
| The outcome | Rejecting a false , the correct call | Failing to reject a false , a miss |
| Probability | ||
| Possible only when | is false | is false |
| Needs a stated alternative value | Yes | Yes |
| Chosen by you | No | No |
| In the baseline test above | 0.5087 against | 0.4913 against |
| Relationship to | Equals when the truth is the null value | None fixed: here 0.05 and 0.4913 |
Wherever the two columns differ they are describing one event from opposite ends, and the two numeric rows differ by nothing more than a subtraction. The rest of the rows are shared outright, because these are one event counted twice. In practice the choice between the words is just whichever one the question asks you to name.
The mix-ups worth killing
Five sentences to stop writing, each with the correction attached.
- "Power is 0.5087, so the null is about 51 percent likely to be false." Power says nothing about whether is false. It assumes a stated alternative value is already the truth and reports how often the test would notice.
- "Alpha and beta add to 1." They do not. In the baseline test they are 0.05 and 0.4913, which sum to 0.5413 and mean nothing. Power and are the complements; and are unrelated numbers about two different states of the world.
- "My p-value was 0.31, so there is a 31 percent chance I made a Type II error." A p-value is computed assuming is true, and under a true a Type II error cannot occur at all. is never read off your data; see what a p-value means.
- "This test has a power of 0.80." Against what? Without the alternative value the sentence carries no information, and the table above is why: the same test ran from 0.0500 to 0.9543 with nothing changed but the assumed truth.
- "We failed to reject, so there is no effect." A failure to reject means the evidence was not convincing, not that is true. A low-power test fails to reject constantly even when an effect is real, which is exactly why "no convincing evidence" and "no effect" are different claims. The full reasoning is in Type I vs Type II error and power.
Power and beta against one stated alternative
A z test uses against at , with known and . Find the power and the probability of a Type II error when the true mean is 107.
Standard error first: .
Put the rejection rule on the sample-mean scale. The upper critical value for is , the four-decimal version of the 1.645 a z table prints, so the test rejects when .
Now switch distributions. If the truth is , sample means center at 107 with the same standard error of 3.
Standardize the cutoff against that true distribution: .
Power is the chance of landing in the rejection region under the truth: .
The Type II error is the complement: .
Check the identity: , which is what requires.
Power is 0.7544 and , both against and against nothing else. Notice that step 2 used no information about the truth and step 4 used nothing else: the cutoff belongs to the test, the two probabilities belong to the test plus one named alternative.
One test, two powers
Keep that test exactly as it was: against , , , , rejecting when . A student writes "the power of this test is". Compute the power against and against , then say what is wrong with the phrase.
The cutoff is built from , , , and only, so it stays at 104.9347 for both, and the standard error stays at 3.
Against : , so power is and .
Against : , so power is and .
Compare the inputs. Nothing about , , , the hypotheses, or the cutoff differs between the two calculations.
The only thing that changed is the alternative value the power was aimed at, and the answer moved from 0.2595 to 0.8466.
0.2595 against and 0.8466 against . There is no single power of this test, so the phrase is incomplete rather than wrong: a power or a becomes a real number only once the alternative value it targets is named. That is also why can never be recovered from your sample, which knows nothing about which alternative you had in mind.
Frequently asked questions
Is power just 1 minus the Type II error rate?
Yes. Power is the probability of correctly rejecting a false null hypothesis and beta is the probability of failing to reject it, so the two always add to 1. The catch is that both are computed against one specific true parameter value, so a power and a beta only pair up when they name the same alternative.
Do alpha and beta add to 1?
No, and they have no fixed relationship at all. In the z test on this page alpha is 0.05 and beta against a true mean of 105 is 0.4913, which sum to 0.5413. Alpha is the error rate when the null is true; beta is the error rate when it is false. Power and beta are the complements, not alpha and beta.
Can power ever be smaller than alpha?
Yes, for a one-sided test aimed the wrong way. The upper-tailed test here has power 0.0005 against a true mean of 95, far below its alpha of 0.05, because it cannot detect an effect in the direction it is not looking. A two-sided test cannot fall below alpha: its power curve bottoms out at alpha when the truth equals the null value.
Can I work out beta from my p-value?
No. A p-value is computed assuming the null hypothesis is true, and a Type II error is only possible when the null is false, so the p-value carries no information about beta. Beta is a property of the test aimed at an alternative value you never get to observe, not something the data report.
How do I raise power without raising the Type I error rate?
Increase the sample size, or reduce variability through a cleaner measurement process or a better design. In the test on this page, moving n from 25 to 100 takes power against a true mean of 105 from 0.5087 to 0.9543 while alpha stays at 0.0500. Raising alpha also raises power, but it buys that by making false alarms more common.