Hypothesis testing
Every term in the four-step test, including the two errors and the ones students most often state backwards.
22 terms
Alternative hypothesisThe alternative hypothesis, written Ha, is the claim about a population parameter that decides which departures from the null count as evidence.Chi-square testA chi-square test compares observed counts of categorical data to the counts expected under a hypothesis, gauging how far the data stray from that model.Decision ruleThe decision rule is the standard you set before collecting data: reject the null hypothesis when the p-value is at or below alpha.Effect sizeEffect size measures how big a difference is, in units you can interpret, separately from whether that difference is statistically significant.Expected countAn expected count is how many observations a category would get if the null hypothesis were exactly true; it is the baseline in a chi-square test.Hypothesis testA hypothesis test uses sample data to weigh a null claim against an alternative, gauging how surprising the data would be if the null claim were true.Null distributionThe null distribution is the sampling distribution of a statistic when the null hypothesis is true; a p-value is an area in its tail.Null hypothesisThe null hypothesis, written H0, is the default claim of no effect or no difference, assumed true unless the sample data give strong evidence against it.One-sided testA one-sided test has an alternative hypothesis using < or >, so only departures from the null in one direction count as evidence.P-valueA p-value is the probability, computed assuming the null hypothesis is true, of getting a result at least as extreme as the one you observed.PowerPower is the probability that a test rejects the null hypothesis when one specific alternative value is the truth, equal to 1 minus the Type II error rate.Practical significancePractical significance asks whether an effect is large enough to matter in context, a judgment no p-value can make for you.Randomization testA randomization test builds its null distribution by randomly reassigning the observed responses to the treatment groups, rather than reading it off a formula.Rejection regionThe rejection region is the set of test statistic values extreme enough to reject the null hypothesis at a chosen significance level alpha.RobustnessA procedure is robust when it still gives roughly correct results even though one of its conditions is mildly violated.Significance levelThe significance level, alpha, is the threshold a p-value is compared to, set before testing as the accepted probability of rejecting a true null hypothesis.Standardized test statisticThe standardized test statistic is the statistic minus the parameter value the null hypothesis claims, divided by the standard error of the statistic.Statistical significanceA result is statistically significant when its p-value is at or below alpha: a result at least this extreme would be unlikely if the null hypothesis were true.Test statisticA test statistic is a single number computed from sample data that measures how far the sample falls from what the null hypothesis predicts.Two-sided testA two-sided test has an alternative hypothesis using a not-equal sign, so a departure from the null in either direction counts as evidence.Type I errorA Type I error is rejecting a true null hypothesis: a false positive, concluding there is an effect when in fact there is none.Type II errorA Type II error is failing to reject a false null hypothesis: a false negative, missing a real effect that is actually present.