Degrees of Freedom vs T-Distribution
Both terms below come up in the same part of the course, and students mix them up. Here is each one defined on its own, side by side, so you can see where they part company.
Degrees of freedom
Confidence intervals
Degrees of freedom count the independent pieces of information left after estimation, and they select which t or chi-square curve a statistic follows.
Degrees of freedom, written , count the independent pieces of information a statistic has left after some are spent estimating other quantities. Make that concrete. Four measurements with a sample mean of 10 must total 40. Pick the first three freely, say 7, 12, and 6, and the fourth is forced: . Three values were free, so .
The count matters because it selects the curve that sets the multiplier. A one-sample interval at 95% uses at degrees of freedom: 2.262 when , 2.064 when , and 1.984 when . Same level, three multipliers, differing only in how much information carries. All approach from above without arriving.
is not a universal rule. A paired uses one fewer than the number of pairs, not of measurements. A chi-square test on an by table uses , which depends on the shape of the table, not the sample size: a 2 by 3 table carries 2 degrees of freedom whether the counts total 60 or 6000.
Two habits cause most of the damage. One is reading the table at row : a sample of 20 has 19 df and , while row 20 gives 2.086. The other is expecting a whole number: Welch's two-sample with , , , and returns . Nothing is counted there: it is the curve that best approximates a statistic whose distribution is not , and that density exists for any positive .
The Welch value is penned in by , so 23.40 must land between 11 and 25, a fast data entry check. A calculator reports the decimal; a printed table forces a whole row, so round down. is 2.069 at 23 df against 2.067 at 23.40, so the interval comes out slightly wide, not falsely narrow.
t-distribution
Random variables and distributions
The t-distribution is a symmetric, bell-shaped curve with heavier tails than the normal, used for inference about a mean when the population SD is unknown.
The -distribution is not one curve but a family, indexed by the degrees of freedom. It is the distribution of (x-bar minus mu, over s divided by root n) when the data come from a normal population. Swapping the fixed (sigma) for the sample standard deviation , which itself changes from sample to sample, is what puts the extra weight in the tails. Every member is symmetric about 0, and the family closes on the standard normal as the degrees of freedom grow.
The numbers make that convergence concrete. For a 95 percent interval the t-table gives at 14 degrees of freedom, at 30, at 100 and at 1000, against for the normal. The gap is 8.6 percent of the critical value at 14 degrees of freedom and about 0.1 percent at 1000.
"The sample is small so use , and large so use " is the wrong rule, and it is the one most students carry in. The trigger is whether is known, not how big is. With 500 observations and a standard deviation estimated from them, the correct model is on 499 degrees of freedom, which happens to sit very close to the normal. Knowing with would put you back on .
The heavier tails change verdicts, not just widths. A statistic of 2.00 read on with 14 degrees of freedom has a two-sided p-value of 0.0653, against the 0.0455 the normal returns for the same 2.00, so at one model rejects and the other does not.
Those tails cover the uncertainty in and nothing else. They do not repair a skewed population or a stray outlier, which is why a procedure still asks you to look at the shape of the sample first. The -distribution enters the course at topic 4.2 of Unit 4, Inference for Quantitative Data: Means.