Type I and Type II Error Practice Problems
By Jude Wallis · Published
This set has 8 problems on Type I and Type II errors. Each asks you to name the error in context, describe its consequence, and say how alpha, sample size, and the size of the true effect move the two error rates and power. Solve each on paper before opening the steps.
AP Statistics: Unit 3 (topics 3.8 Potential Errors When Performing Tests). These problems cover Unit 3 topic 3.8 (Potential Errors When Performing Tests) in the Fall 2026 AP Statistics course, where you identify Type I and Type II errors in context, describe their consequences, and explain what moves the error rates and power.
What these problems build
These 8 problems build one workflow: work out what a test's decision could be getting wrong, name it, and say what it costs someone in the scenario. A Type I error is rejecting when is true, and its probability is the significance level that you choose before the test. A Type II error is failing to reject when is false, and its probability is , which you do not set directly. Power is , the chance of correctly rejecting a false null.
| Your decision | is true | is false |
|---|---|---|
| Reject | Type I error, probability | Correct decision, probability power |
| Fail to reject | Correct decision | Type II error, probability |
Read the table by row and one fact falls out that these problems lean on constantly: only one error is even possible for a given completed test. If you rejected, the mistake on the table is a Type I error. If you failed to reject, it is a Type II error. Naming both for a single decision is the most common way students lose the point. For the concepts behind the table, read type I vs type II errors.
How alpha, sample size, and effect size move the two rates
Half of these problems ask you to predict which error a change makes more likely. Take one change at a time and hold the rest fixed.
- **Lower **, say from 0.05 to 0.01. The Type I error rate falls to 0.01, but rejecting now takes stronger evidence, so rises and power falls. A Type II error becomes more likely.
- **Raise .** The mirror image: more power and fewer Type II errors, at a higher rate of false alarms.
- **Raise the sample size .** The standard error shrinks and the test separates the null value from the truth more sharply, so power rises while stays exactly where you set it. This is the lever you actually control that improves one error rate without worsening the other.
- A true parameter farther from the null. Bigger real effects are easier to catch, so power rises and falls. You do not control the truth, but problems compare two alternatives, as problem 7 does.
- Less variable measurement. A smaller standard error acts like a larger : more power, same .
One arithmetic fact worth carrying in: quadrupling halves the standard error and therefore doubles the test statistic when the sample proportion stays the same. Problem 6 uses exactly that to turn a near miss into a decisive rejection.
Reading the table and phrasing the conclusion
Every test you actually carry out in this set, in problems 5, 6, and 8, is a one-sample z-test for a proportion, so those p-values come from the z-table, which gives the area to the left of a value to four decimal places. That is a direct read. If one of these scenarios called for a t procedure instead, the exam's t-table lists only a handful of tail probabilities per row, so you would bound the p-value between two columns rather than report a single number; the t-test practice set works that way throughout. Setting up the proportion test itself is covered in one proportion vs two proportion z test.
When you name an error, three things earn the credit and a bare textbook definition earns none:
- Say which error it is, Type I or Type II.
- Describe it using the actual claim rather than the symbols. "Concluding more than 2% of sensors are defective when the true rate is 2%" beats "rejecting a true null hypothesis."
- State the real-world consequence for someone in the scenario.
For how the p-value drives the decision that the error attaches to, see what does a p-value mean and p-value vs alpha. More sets are on the practice page.
Frequently asked questions
How do I know which error a completed test could be?
Look at the decision you made, not at the truth you cannot see. If you rejected the null hypothesis, the only possible mistake is a Type I error. If you failed to reject it, the only possible mistake is a Type II error. Naming both for one decision loses the point.
Does lowering alpha reduce both error rates?
No. Lowering alpha lowers the Type I error rate but demands stronger evidence to reject, so beta rises and power falls, making a Type II error more likely. Increasing the sample size, or measuring more precisely, lowers the Type II error rate while leaving alpha exactly where you set it.
How much detail does an error description need on the exam?
Three parts: name the error as Type I or Type II, restate it using the actual claim in the scenario rather than the symbols, and give the real-world consequence for someone involved. A correct textbook definition with no context earns nothing.
Do I need a formula for statistical power?
No. The course expects you to find error probabilities from given values, from a simulation, or from a sampling distribution the problem describes, which is what problems 3 and 7 in this set do. There is no power formula to memorize.
Problem 1
A city water utility tests whether more than 10% of homes in a district have lead above the federal action level. Let be the true proportion of homes in that district above the action level. The utility tests against at . (a) Describe a Type I error in context and give its consequence. (b) Describe a Type II error in context and give its consequence. (c) Give the probability of a Type I error. (d) Which error is more serious here?
Show the worked solution
Read what claims. says the district sits right at the action level, so there is no evidence of a lead problem reaching beyond 10% of homes.
(a) A Type I error rejects a true . Here that means concluding more than 10% of homes exceed the action level when in truth only 10% or fewer do. Consequence: the utility tears up and replaces service lines across a district that did not need the work, spending money and disrupting residents for nothing.
(b) A Type II error fails to reject a false . Here that means finding insufficient evidence when more than 10% of homes really are above the action level. Consequence: no replacement program starts and residents keep drinking water with elevated lead.
(c) If is true, the probability of a Type I error is the significance level the utility set in advance, .
(d) Compare the two consequences rather than the two labels. A Type I error costs money and disruption. A Type II error leaves a population exposed to lead, which does irreversible harm to children. The Type II error is the more serious one here, which argues for a larger or a larger sample so the test has more power.
(a) Type I: deciding more than 10% of homes exceed the action level when they do not, causing needless pipe replacement. (b) Type II: missing a real problem, leaving residents exposed to lead. (c) . (d) The Type II error, because continued lead exposure outweighs unnecessary construction.
Problem 2
A nutrition lab tests whether a redesigned food label raises the proportion of shoppers who check calories before buying. It tests against at with shoppers. Against the specific alternative , the test has power . (a) Find for that alternative. (b) Interpret the power in context. (c) Taking one change at a time, state what happens to power and which error becomes more likely: (i) raising from 200 to 500, (ii) lowering from 0.05 to 0.01, (iii) the true proportion being 0.34 instead of 0.38.
Show the worked solution
(a) Power and are complements against a fixed alternative, so .
(b) Interpret in context. If the redesigned label truly raises the proportion who check calories to 0.38, this test detects that increase and rejects about 78% of the time. The other 22% of the time it misses a real improvement, which is a Type II error.
(c)(i) Raising to 500 shrinks the standard error, so power rises above 0.78 and falls below 0.22. A Type II error becomes less likely. The Type I error rate is untouched at , because you set it directly.
(c)(ii) Lowering to 0.01 cuts the Type I error rate to 0.01 but demands stronger evidence to reject, so power falls below 0.78 and rises. A Type II error becomes more likely. This is the genuine tradeoff between the two errors.
(c)(iii) A true proportion of 0.34 sits closer to the null value of 0.30 than 0.38 does, so the effect is harder to detect. Power falls below 0.78 and rises, making a Type II error more likely, while stays 0.05.
(a) . (b) If the true proportion is 0.38, the test catches the increase about 78% of the time and misses it 22% of the time. (c)(i) power up, Type II less likely, still 0.05; (ii) power down, Type II more likely, Type I rate down to 0.01; (iii) power down, Type II more likely, unchanged.
Problem 3
A statistics class studies a test's performance by simulating 500 random samples from a population in which the null hypothesis is false, so a real effect exists. Their test rejects in 340 of the 500 simulated tests. (a) Estimate the power of the test. (b) Estimate . (c) If the class ran 75 more studies under these same conditions, about how many would end in a Type II error? (d) If they redid the simulation at instead of , would that count go up or down, and what is the cost?
Show the worked solution
(a) Power is the probability of correctly rejecting a false , so estimate it with the fraction of simulations that rejected: .
(b) The remaining simulations failed to reject a null that was false, which is exactly a Type II error. That happened times, so .
Check the complement. , as required by .
(c) Each future study under the same conditions ends in a Type II error with probability about 0.32, so the expected count in 75 studies is .
(d) A larger makes rejection easier, so power rises and falls. The expected number of Type II errors would go down, below 24. The cost is that the Type I error rate doubles from 0.05 to 0.10, so studies of populations with no real effect would raise false alarms twice as often.
(a) power . (b) . (c) About Type II errors. (d) Down, because a larger raises power, but the Type I error rate rises from 0.05 to 0.10.
Problem 4
A research group runs 40 independent hypothesis tests, each at , and suppose every one of the 40 null hypotheses is actually true. (a) How many Type I errors should the group expect? (b) Find the probability that at least one of the 40 tests produces a Type I error. (c) How many Type II errors should the group expect?
Show the worked solution
(a) When is true, a test produces a Type I error with probability . Counting errors across 40 independent tests is binomial with and , so the expected count is .
(b) Go through the complement, since "at least one" is awkward to count directly. A single test avoids a Type I error with probability , and the tests are independent, so all 40 avoid it with probability .
(b) Evaluate. , so .
(c) A Type II error requires failing to reject a null that is false. Every null here is true, so no test in this batch can commit one. The expected number of Type II errors is 0.
Interpret. Running many tests at makes at least one false alarm the likely outcome, here about 87%, even when nothing real is happening, which is why one significant result pulled from a pile of tests is weak evidence on its own.
(a) expected Type I errors. (b) . (c) Zero, because a Type II error requires a false null and all 40 nulls are true.
Problem 5
A delivery service advertises that 80% of its orders arrive within two days. A consumer group takes a random sample of 150 recent orders and finds that 108 arrived within two days. (a) Carry out a test at of whether the true proportion arriving within two days is less than 0.80. (b) Name the error this decision could be, describe it in context, and give its probability if is true.
Show the worked solution
State. Let be the true proportion of all the service's orders that arrive within two days. versus , with .
Plan. One-sample z-test for a proportion. Random: the 150 orders are a random sample. 10%: 150 is less than 10% of all orders the service handles. Large counts: and , both at least 10.
Do. The sample proportion is . The standard error uses the null value: .
Do. . The test is left-tailed, so the p-value is the area below , which the z-table gives directly as .
Conclude. Since the p-value is less than , reject . There is convincing evidence that the true proportion of the service's orders arriving within two days is less than 0.80.
(b) The decision was to reject, so the only mistake still available is a Type I error: concluding that fewer than 80% of orders arrive within two days when in truth exactly 80% do. Consequence: the group publishes a false accusation and the service takes reputational damage it did not earn. If is true, the probability of this error is .
(a) , , p-value ; reject at , so there is convincing evidence fewer than 80% of orders arrive within two days. (b) A wrong rejection here is a Type I error, probability 0.05: falsely accusing a service that does meet its 80% claim.
Problem 6
A nurse at a high school with 4,000 students suspects that more than 25% of them skip breakfast. In a random sample of 80 students, 26 report skipping breakfast. (a) Carry out a test at . (b) Name the error this decision could be and describe it in context. (c) The nurse repeats the study with a random sample of 320 students and gets the same sample proportion, 104 of 320. Recompute the test statistic and p-value, then explain what the comparison shows about sample size and power.
Show the worked solution
State. Let be the true proportion of students at the school who skip breakfast. versus , with .
Plan. One-sample z-test for a proportion. Random: the 80 students are a random sample. 10%: 80 is 2% of the school's 4,000 students, well under 10%. Large counts: and , both at least 10.
Do. and , so . The test is right-tailed, so the p-value is from the z-table.
Conclude. Since , fail to reject . There is not convincing evidence that more than 25% of students at the school skip breakfast.
(b) The decision was to fail to reject, so the only mistake still available is a Type II error: finding insufficient evidence when more than 25% of students really do skip breakfast. Consequence: the school declines to fund a breakfast program that students genuinely need.
(c) Recheck the conditions before recomputing, since the sample got four times bigger. Random: still a random sample. 10%: 320 is 8% of the school's 4,000 students, still under 10%. Large counts: and , both at least 10.
(c) With the sample proportion is unchanged at , but . Quadrupling halves the standard error, so , exactly double the earlier statistic. The p-value is now , and since you reject .
(c) Interpret. The same observed gap of 7.5 percentage points went from not significant at to decisive, with no change to . A larger sample raises power, which lowers the Type II error rate while leaving the Type I error rate at 0.05. That is why the fix for a near-miss p-value is more data, not a looser .
(a) , , p-value ; fail to reject at . (b) A Type II error: missing a real skip rate above 25% and leaving a needed breakfast program unfunded. (c) With , and the p-value is , so reject. Quadrupling halves , doubles , and raises power with fixed.
Problem 7
A polling firm tests against using a random sample of voters, and its decision rule is to reject whenever . (a) If the true proportion is , find the power of this test and the probability of a Type II error. (b) Repeat for a true proportion of . (c) Explain what the comparison shows. In each part the sampling distribution of is approximately Normal.
Show the worked solution
(a) Describe the sampling distribution under the stated truth, not under the null. With and , the counts and are both at least 10, so is approximately Normal with mean and standard deviation .
(a) Power is the probability that this test rejects, and the rule rejects when . Standardize the cutoff: .
(a) The z-table gives the area below as , so . The Type II error probability is the complement, .
(b) Now take . The counts and are both at least 10, and the standard deviation is .
(b) Standardize the same cutoff: . The area below is , so and .
(c) Interpret. The sample size, the cutoff, and therefore the Type I error rate are identical in both parts; only the true proportion moved. A truth of 0.60 sits closer to the null value of 0.50 than 0.65 does, so power drops from 0.929 to 0.659 and the Type II error rate rises from 0.071 to 0.341. Real effects close to the null are the ones tests miss.
(a) , so power and . (b) , so power and . (c) With and held fixed, a true proportion closer to the null gives lower power and a larger Type II error rate.
Problem 8
A manufacturer's contract requires that no more than 2% of the pressure sensors in a production run be defective. An independent inspector draws a random sample of 500 sensors from a large run and finds 16 defective. (a) State the hypotheses and check the conditions. (b) Find the test statistic and p-value. (c) State a conclusion at , then name the error that decision could be and give its consequence. (d) The manufacturer argues the test should use . Give the conclusion at that level, name the error that decision could be, and say which party each significance level protects.
Show the worked solution
(a) State. Let be the true proportion of defective sensors in the run. versus , because the inspector is checking whether the run exceeds the contract limit.
(a) Plan. One-sample z-test for a proportion. Random: the 500 sensors are a random sample from the run. 10%: 500 is less than 10% of a large production run. Large counts: and , both at least 10, with sitting exactly on the boundary.
(b) Do. and .
(b) . The test is right-tailed, so the p-value is from the z-table.
(c) At , since , reject . There is convincing evidence that more than 2% of the sensors in this run are defective. Because the decision was to reject, the only mistake available is a Type I error: declaring the run out of specification when it truly meets the 2% limit. Consequence: the manufacturer scraps or reworks an acceptable run and absorbs the cost.
(d) At , since , fail to reject . There is not convincing evidence that more than 2% of the sensors are defective. Because the decision was to fail to reject, the only mistake available is a Type II error: passing a run whose true defect rate really does exceed 2%. Consequence: defective pressure sensors ship to customers.
(d) Compare. The same data flips the decision because the cutoff moved, not because the evidence changed. The larger carries the higher Type I error rate but also the higher power, so it protects the buyer against a bad run slipping through. The smaller makes rejection harder, which protects the manufacturer against scrapping a run that was fine. Which level is appropriate depends on whether a shipped defective sensor or a wasted good run is the costlier failure.
(a) against , conditions met with exactly at the boundary. (b) , , p-value . (c) Reject at ; a wrong rejection is a Type I error, scrapping a run that met the 2% limit. (d) Fail to reject at ; a wrong failure to reject is a Type II error, shipping defective sensors. The larger protects the buyer, the smaller protects the manufacturer.