Classifying variables practice problems

By Jude Wallis · Published

This set has 8 problems on AP Statistics topic 1.2. Every one runs the same two steps: decide categorical or quantitative with the arithmetic test, then, only for quantitative variables, decide discrete or continuous. Solve each on paper before opening the solution.

AP Statistics: Unit 1 (topics 1.2 Variables, 1.4 Graphical Representations for One Categorical Variable, 1.5 Graphical Representations for One Quantitative Variable). These problems drill Unit 1 topic 1.2 (Variables) of the Fall 2026 AP Statistics course, where Unit 1 carries 20 to 30% of the multiple-choice section. Problems 6 and 7 also reach topics 1.4 and 1.5, since the type of a variable is what decides whether its graph is a bar graph or a histogram.

The two-step test these problems drill

Classification is two decisions in a fixed order, and doing them out of order is what turns a small slip into a wrong classification.

  1. Categorical or quantitative? Ask whether arithmetic on the values means anything. If adding or averaging the values gives a number that measures something, the variable is quantitative. If the values only say which group an individual is in, it is categorical (the course also calls this qualitative).
  2. Only if quantitative, discrete or continuous? Ask whether you could list the possible values one at a time. A list, finite or endless, means discrete. Values that fill an interval, so that between any two there is always another, mean continuous.

Step 2 never applies to a categorical variable, so "discrete categorical" is not an answer. For the definitions behind this, read categorical vs quantitative variables and discrete vs continuous variables. These 8 problems are drills on top of those two pages, not a second explanation of them. More sets are on the practice page.

The five cases that actually cost points

Almost every miss on this topic is one of five things, and every problem in this set except problem 6 is built on at least one of them.

  • Digits that are labels. A zip code, an area code, a student ID, and a jersey number are written with digits and measure nothing, so they are categorical. Reissue the numbers and the average moves to wherever the new labels put it, while nobody's data changed. See nominal variable.
  • A count is quantitative and discrete. Counting is not the same as categorizing. The number of pets in a home is a counted quantity you can average, and its possible values are 0, 1, 2, and so on, so it is quantitative and discrete.
  • Precision is a property of the instrument, not the variable. Age written in whole years is still continuous, because the quantity behind it is elapsed time. Judge by what is being measured, not by how coarsely it was recorded. See continuous variable.
  • Decimals do not prove continuity, and whole numbers do not prove discreteness. A discrete variable can have a decimal mean, and a variable whose recorded values are decimals can still have a listable set of possible values.
  • Ordered categories are still categorical. Small, medium, large, a letter grade, and a 1-to-5 satisfaction rating all carry a real order and no measured distance between neighbors. Problem 4 works through what that costs, and it is the one case on this page with a genuinely arguable edge: see ordinal variable.

Naming the variable a study actually measured

Problems 5, 6, and 8 ask a harder question than "which type is this?" They ask what was measured and on whom. Two habits carry those questions.

First, separate the observational unit from the variable. The unit is who or what a datum came from, one row of the data set. The variable is the characteristic recorded about it, one column. A summary such as "73% were vaccinated" is neither: it is a number computed from a column. Change the unit and the same subject matter produces a different variable, which is exactly the trap in problem 5.

Second, notice that one characteristic can be recorded as more than one variable. Screen time measured in minutes is quantitative; the same students answering "more than two hours, yes or no" produce a categorical variable about the same behavior. Which one you collected decides which summary, which graph, and which later unit of the course applies, and that conversion runs one way only. Problem 6 makes that concrete.

Frequently asked questions

Can a categorical variable be discrete or continuous?

No. Discrete and continuous are subtypes of quantitative variables only, so the second step of the test never runs on a categorical variable. If you have written "discrete categorical" on an answer, go back to step 1: either the variable is categorical and the classification stops there, or it is quantitative and the word categorical is wrong.

Is a variable made of numbers always quantitative?

No. Zip codes, area codes, student IDs, jersey numbers, and the codes a survey assigns to answer choices are digits used as names, and they are categorical. The test in problem 2 is the practical one: relabel the categories in an order-preserving or arbitrary way, and if a summary moves while nobody's data changed, that summary was never measuring anything.

Are counts categorical, since they come from sorting things into groups?

The count is quantitative and discrete; the thing being counted may be categorical. Eight households reporting 0, 1, and 2 pets give a quantitative variable measured on each household. Tallying 120 students by favorite lunch gives frequencies that are also numbers, but the variable there is favorite lunch, which is categorical. Ask which characteristic varies from row to row.

If a value is recorded to the nearest whole number, is the variable discrete?

Not by itself. Every measurement is rounded somewhere, so reading rounding as evidence of discreteness would make every variable discrete. Age in whole years, height to the nearest centimeter, and a time to the nearest tenth of a second are all continuous, because the quantity behind the record fills an interval. Judge the quantity, not the precision of the instrument.

Can I report the mean of a 1-to-5 satisfaction rating?

It is common practice, and problem 4 shows what it costs: two codings that respect the same order gave improvements of 0.275 and 1.275 from the same two sets of 40 ratings. The median category and the share rating 4 or 5 do not move under recoding, so lead with those. If you do report a mean of the codes, say out loud that you are treating the five steps as equally sized, which nothing in the survey measured.

Problem 1

Classify each variable as categorical or quantitative. For every quantitative variable, also say whether it is discrete or continuous.

(a) The zip code of a customer's home address. (b) The number of siblings a student has. (c) A patient's height in centimeters. (d) A donor's blood type. (e) The time in seconds a runner takes to finish 400 meters. (f) The number of items in a shopper's basket at checkout. (g) A student's letter grade in a course (A, B, C, D, or F). (h) The rainfall recorded at a weather station in one day, in millimeters.

Show the worked solution
  1. Run step 1 on all eight: would adding or averaging the values give a number that measures something? Run step 2 only on the ones that survive step 1.

  2. (a) Zip code. An average zip code points to no location, and 90210 is not more of anything than 10001. The digits are an address label, so this is categorical and step 2 does not apply.

  3. (b) Number of siblings. An average number of siblings is a real quantity, so it is quantitative. The possible values are 0, 1, 2, 3, and onward, listable with nothing possible between 2 and 3, so it is discrete.

  4. (c) Height. An average height is meaningful, so it is quantitative. Between any two heights another height is possible, so it is continuous. A ruler reading to the nearest centimeter does not change that.

  5. (d) Blood type. O, A, B, and AB are names with no order and no average, so this is categorical.

  6. (e) Time to run 400 meters. An average time is meaningful, so quantitative, and elapsed time fills an interval, so continuous.

  7. (f) Number of items in a basket. A count you can average, so quantitative, and the possible values are 0, 1, 2, and onward, so discrete.

  8. (g) Letter grade. A, B, C, D, and F are category labels. They carry an order, which is more than blood type has, but there is no measured distance between a B and a C, so this is categorical.

  9. (h) Daily rainfall. Depth of water is measured and averaged, so quantitative, and it fills an interval, so continuous. Many days record 0.0 mm, and a pile of zeros is a fact about the weather, not evidence that the variable is discrete.

Categorical: (a) zip code, (d) blood type, (g) letter grade. Quantitative and discrete: (b) number of siblings, (f) number of items in a basket. Quantitative and continuous: (c) height, (e) 400 meter time, (h) daily rainfall.

Problem 2

A youth soccer league records two numbers for each of the six players on a team: the jersey number and the goals scored this season.

Player123456
Jersey number479121822
Goals scored011235

(a) Classify each of the two variables. (b) Compute the mean of each row and say which one is a legitimate summary. (c) Over the winter the league reissues jerseys as 1 through 6, with no change to the roster or to anyone's goals. Recompute both means. (d) What does the comparison in (c) prove?

Show the worked solution
  1. (a) Jersey number is a name written with digits: it identifies a player and measures no quantity, so it is categorical. Goals scored counts a quantity you can add, so it is quantitative, and its possible values are 0, 1, 2, and onward, so it is discrete.

  2. (b) Jersey numbers: the sum is 4+7+9+12+18+22=724 + 7 + 9 + 12 + 18 + 22 = 72, so the mean is 726=12\frac{72}{6} = 12. Goals: the sum is 0+1+1+2+3+5=120 + 1 + 1 + 2 + 3 + 5 = 12, so the mean is 126=2\frac{12}{6} = 2.

  3. (b) Only the second is a summary of anything. "This team averages 2 goals per player" describes the season. "This team averages jersey number 12" describes nothing, and the fact that one player does wear 12 is a coincidence of these six labels: move that jersey to 13 and the mean becomes 73612.17\frac{73}{6} \approx 12.17, a number nobody wears, without changing anything about the team.

  4. (c) After renumbering, the jersey sum is 1+2+3+4+5+6=211 + 2 + 3 + 4 + 5 + 6 = 21, so the mean jersey number is 216=3.5\frac{21}{6} = 3.5. The goals did not change, so the mean goals is still 126=2\frac{12}{6} = 2.

  5. (d) The mean jersey number moved from 12 to 3.5 while not one fact about the six players changed, and the mean goals stayed at exactly 2. That is the operational meaning of "arithmetic on the values is meaningless": a summary that moves when someone relabels the categories was never measuring the players.

  6. One more distinction worth naming. Jersey number here is functioning as an identifier, the way the Player row does, so it labels the rows rather than describing them. Identifiers are categorical when they are treated as variables at all, and they are usually not graphed or summarized.

(a) Jersey number is categorical; goals scored is quantitative and discrete. (b) Mean jersey number 726=12\frac{72}{6} = 12, mean goals 126=2\frac{12}{6} = 2; only the mean goals summarizes anything. (c) After renumbering, mean jersey number 216=3.5\frac{21}{6} = 3.5 and mean goals still 2. (d) Relabeling the categories moved the jersey mean by 8.5 without changing any player, which is why the jersey mean is not a summary of the data.

Problem 3

All six variables below are quantitative. Label each one discrete or continuous, and give the test that produced your label.

(a) A person's age, recorded in whole years on a form. (b) The proportion of 20 free throws a player makes. (c) The volume of juice in a carton, in milliliters. (d) The number of defective bulbs in a box of 24. (e) The price of a sandwich, in dollars and cents. (f) The length of a leaf, recorded to the nearest whole centimeter.

Show the worked solution
  1. The test is not what the recorded numbers look like. Ask what set of values the quantity could take, then ask whether that set can be listed one value at a time.

  2. (a) Age. The quantity behind the record is elapsed time, which fills an interval: between 15 years and 16 years every age in between is possible and real. The form rounds it, and rounding is the instrument. Age is continuous.

  3. (b) Proportion of 20 free throws made. The player makes some whole number kk of shots from 0 to 20, so the proportion can only be k20\frac{k}{20}: the 21 values 0,0.05,0.10,,1.000, 0.05, 0.10, \ldots, 1.00. Nothing at 0.07 is possible. Decimals throughout and still discrete, because those 21 values can be listed one at a time with nothing possible in between. Shortness is not the test, listability is: a list of 21 values and a list that never stops are both lists.

  4. (c) Volume of juice. Volume fills an interval, so between any two fill levels another is possible. Continuous.

  5. (d) Number of defective bulbs. The possible values are the 25 whole numbers 0 through 24, a finite list. Finite counts as listable, so this is discrete. Discrete does not require the list to be endless.

  6. (e) Price. A price is a whole number of cents, so its possible values are countable and strictly it is discrete. It is routinely modeled as continuous, because a cent is tiny next to typical prices, and either answer is defensible as long as you say which one you are using and why. Note that (b) and (e) are not the same situation for that modeling choice, though both are discrete under the listability test: 21 values spaced 0.05 apart are nowhere near a continuum, while the possible prices under 20 dollars run into the thousands, spaced a cent apart, which is close enough to one that treating them as continuous costs almost nothing.

  7. (f) Leaf length. Length fills an interval, and recording to the nearest centimeter is a choice about the ruler. Continuous.

  8. Read (a) and (b) together, because they break the two halves of the same false rule. "Whole numbers mean discrete" fails on age, which is written as an integer every time and is continuous. "Decimals mean continuous" fails on the free throw proportion, which is decimal every time and is discrete.

Discrete: (b) the free throw proportion, which can only be one of the 21 values k20\frac{k}{20}, and (d) the defect count, a finite list 0 through 24. Continuous: (a) age, (c) volume, (f) leaf length, since each measures a quantity that fills an interval and is merely recorded coarsely. (e) Price is strictly discrete in whole cents and is usually modeled as continuous; say which convention you are using.

Problem 4

A cafe asks 40 customers to rate service on a 5-point scale, where 1 is poor and 5 is excellent, and compares this year with last year. Both samples have 40 customers.

Rating12345
This year6451015
Last year2812117

(a) Classify the rating variable. (b) The manager treats the ratings as the numbers 1 through 5 and reports that mean satisfaction rose from 3.325 to 3.60, "an improvement of 0.275 points." Verify the arithmetic. (c) Recompute both means using the codes 1, 2, 3, 4, 10, which put the categories in the same order, and say what happens to the size of the improvement. (d) Give two summaries that do not depend on the coding, and compute them for both years. (e) What should the manager report?

Show the worked solution
  1. (a) The five ratings are category labels with a real order and no measured distance between neighbors: nothing in the data says the step from 3 to 4 is the same size as the step from 4 to 5. So the rating is a categorical variable with ordered categories. Wider statistics calls this an ordinal variable; the AP course sorts every variable into categorical or quantitative, and this one lands on categorical.

  2. (b) Check both counts first: 6+4+5+10+15=406 + 4 + 5 + 10 + 15 = 40 this year and 2+8+12+11+7=402 + 8 + 12 + 11 + 7 = 40 last year.

  3. (b) This year, using codes 1 to 5: 6(1)+4(2)+5(3)+10(4)+15(5)=6+8+15+40+75=1446(1) + 4(2) + 5(3) + 10(4) + 15(5) = 6 + 8 + 15 + 40 + 75 = 144, so the mean code is 14440=3.60\frac{144}{40} = 3.60. Last year: 2(1)+8(2)+12(3)+11(4)+7(5)=2+16+36+44+35=1332(1) + 8(2) + 12(3) + 11(4) + 7(5) = 2 + 16 + 36 + 44 + 35 = 133, so the mean code is 13340=3.325\frac{133}{40} = 3.325. The difference is 3.603.325=0.2753.60 - 3.325 = 0.275. The arithmetic is right.

  4. (c) Now use the codes 1, 2, 3, 4, 10, which keep the categories in exactly the same order. This year: 6(1)+4(2)+5(3)+10(4)+15(10)=6+8+15+40+150=2196(1) + 4(2) + 5(3) + 10(4) + 15(10) = 6 + 8 + 15 + 40 + 150 = 219, so the mean is 21940=5.475\frac{219}{40} = 5.475. Last year: 2(1)+8(2)+12(3)+11(4)+7(10)=2+16+36+44+70=1682(1) + 8(2) + 12(3) + 11(4) + 7(10) = 2 + 16 + 36 + 44 + 70 = 168, so the mean is 16840=4.20\frac{168}{40} = 4.20. The improvement is now 5.4754.20=1.2755.475 - 4.20 = 1.275.

  5. (c) Nobody's rating changed. The improvement is 0.275 under one order-preserving coding and 1.275 under another, so "0.275 points" is a fact about the numbers someone assigned to the labels, not about the service. The direction happened to survive here; it is not guaranteed to, and a coding change can reverse a comparison outright.

  6. (d) The median category. With 40 ratings the median sits between the 20th and 21st values in order. This year the cumulative counts are 6, 10, 15, 25, 40, so positions 16 through 25 are all rating 4 and the median category is 4. Last year the cumulative counts are 2, 10, 22, 33, 40, so positions 11 through 22 are all rating 3 and the median category is 3. Both are read off the counts and the ordering alone, so no order-preserving recoding can move either one.

  7. (d) The share rating 4 or 5. This year 10+1540=2540=0.625\frac{10 + 15}{40} = \frac{25}{40} = 0.625, and last year 11+740=1840=0.45\frac{11 + 7}{40} = \frac{18}{40} = 0.45. That is a rise of 0.6250.45=0.1750.625 - 0.45 = 0.175, or 17.5 percentage points. This uses only counts, so it too is immune to recoding.

  8. (e) Report the median category rising from 3 to 4 and the share rating 4 or 5 rising from 45% to 62.5%. Both are true of the customers rather than of the code sheet. Averaging ordinal codes is common practice, and the honest version of it states the assumption out loud: that the five steps are being treated as equal in size, which nothing in the survey measured. Display the counts with a bar graph, keeping the bars in rating order, not a histogram.

  9. (e) One caution on scope: these are two samples of 40 customers, and nothing here says how they were selected, so neither number is being offered as an estimate for all of the cafe's customers.

(a) Categorical with ordered categories. (b) Correct: 14440=3.60\frac{144}{40} = 3.60 against 13340=3.325\frac{133}{40} = 3.325, a gap of 0.275. (c) Under the order-preserving codes 1, 2, 3, 4, 10 the means are 21940=5.475\frac{219}{40} = 5.475 and 16840=4.20\frac{168}{40} = 4.20, an improvement of 1.275, so the size of the gain is a property of the coding. (d) Median category 4 this year against 3 last year, and the share rating 4 or 5 rose from 1840=0.45\frac{18}{40} = 0.45 to 2540=0.625\frac{25}{40} = 0.625. (e) Report the median category and the 4-or-5 share, and state the equal-spacing assumption if a mean of the codes is reported at all.

Problem 5

A veterinary clinic pulls its records for last year. For each of the 800 dogs seen it has the breed, the weight in kilograms at that dog's first visit, the number of visits during the year, and whether the dog was vaccinated for rabies (yes or no). The clinic's newsletter says: "Our dogs average 18.4 kg, and 584 of them, or 73%, are vaccinated."

(a) Name the observational unit and give nn. (b) Classify all four variables. (c) The newsletter's 73% is not a variable. Say what it is and which variable it came from. (d) Describe exactly what the 18.4 kg figure measured, and on whom. (e) A student says the number of visits is continuous "because a dog could come in any number of times." Correct the reasoning.

Show the worked solution
  1. (a) One row of this data set is one dog, so the observational unit is a dog seen at this clinic last year, and n=800n = 800. It is not a visit: a dog that came four times still contributes one row, and its four visits are a value inside that row.

  2. (b) Breed is a group label with no order and no average, so categorical. Weight in kilograms is a measured quantity that fills an interval, so quantitative and continuous. Number of visits is a count you can average with possible values 1, 2, 3, and onward, so quantitative and discrete. Rabies vaccination is a yes-or-no label, so categorical.

  3. (c) The 73% is a summary statistic computed from the vaccination column: 584800=0.73\frac{584}{800} = 0.73. A variable is a characteristic that varies from dog to dog, and 73% varies from nothing, because there is only one of it for this data set. The variable is "vaccinated for rabies, yes or no", and 73% is the proportion of one of its two categories.

  4. (c) The same number could be a variable in a different data set. Record the percent vaccinated at each of 60 clinics and the observational unit becomes a clinic, n=60n = 60, and percent vaccinated is then a quantitative variable measured on each one. What a number is depends on the unit it was collected from.

  5. (d) The 18.4 kg is the mean of the weight column, and the weight column is not "a dog's weight" in general. It is the weight recorded at that dog's first visit of the year, one value per dog. It says nothing about how a dog's weight changed across the year, and it describes only the 800 dogs this clinic saw, not dogs in the city or dogs of any particular breed.

  6. (e) The student has used the wrong test. "Could be any number" is about how many values there are, and discrete variables can have endlessly many values: 1, 2, 3, and onward never stops. The test is whether the possible values can be listed one at a time with nothing possible between neighbors. A dog cannot make 2.6 visits, so the values are listable and the variable is discrete. Continuous would require that between any two possible values another one is possible, which is true of the weight column and false of the visit count.

(a) The unit is one dog seen last year, n=800n = 800. (b) Breed categorical; weight quantitative and continuous; number of visits quantitative and discrete; vaccination status categorical. (c) It is a summary statistic, 584800=0.73\frac{584}{800} = 0.73, the proportion in the yes category of the vaccination variable, not a variable itself. (d) It is the mean of each dog's weight at its first visit of the year, for the 800 dogs this clinic saw. (e) Having endlessly many possible values does not make a variable continuous; visits are listable as 1, 2, 3, and onward with nothing in between, so the variable is discrete.

Problem 6

A school nurse studies screen time. The same 300 students answer two questions on one form. Question A asks for the number of minutes they spent on a screen yesterday. Question B asks "did you spend more than 2 hours on a screen yesterday? yes or no." In the results, 189 students reported more than 120 minutes on question A, while 216 students answered yes to question B.

(a) Classify the variable each question produces. (b) Give the summary and the graph that fit each. (c) Can the answers to question A be converted into the answers to question B? Can it be done the other way? (d) Both questions were answered by the same 300 students, so what does the difference between 189 and 216 tell you?

Show the worked solution
  1. (a) Question A records minutes, an amount you can average, so it produces a quantitative variable, and elapsed time fills an interval, so it is continuous. Question B records one of two labels, so it produces a categorical variable, with two categories.

  2. (b) For A, report a center and a spread, such as a mean and a standard deviation or a median and an interquartile range, and graph it with a dotplot, stemplot, histogram, or boxplot, because those place values on a number line. For B, report counts and a proportion, and graph it with a bar graph or a pie chart. Putting B's two categories on a histogram would claim its axis was a number line; see histograms vs bar graphs.

  3. (c) A converts to B. Two hours is 120 minutes, so go down the minutes column, mark every student above 120 as a yes, and you have a categorical column: 189300=0.63\frac{189}{300} = 0.63. The conversion works because the categorical variable is defined by a cut point on the quantitative one.

  4. (c) B does not convert to A. A yes tells you a student was above 120 minutes and nothing more, so no mean, median, or standard deviation of minutes can be recovered from it. Collapsing a quantitative variable to a category throws away information permanently, which is the practical reason to collect the finer measurement when you can.

  5. (d) The gap is 216189=27216 - 189 = 27 students, or 27300=0.09\frac{27}{300} = 0.09 of the group. Since the same 300 students answered both questions, this is not sampling variability and not two different groups. At least 27 students gave a pair of answers that contradict each other, and possibly more, since a student who said yes while reporting 90 minutes and a student who said no while reporting 150 minutes partly cancel in that difference of 27. The likely reason is that the two questions do not mean the same thing to a student: "more than 2 hours" invites a rough judgment, while a minute count invites an estimate the student has to construct.

  6. (d) State what it does not tell you. Nothing here says which of the two numbers is closer to the truth. It says the wording of the question is part of what got measured, so the variable is defined by the instrument that produced it and not only by the behavior it is aimed at.

(a) Question A produces a quantitative continuous variable; question B produces a categorical variable with two categories. (b) A: center and spread, with a dotplot, stemplot, histogram, or boxplot. B: counts and a proportion, with a bar graph or pie chart. (c) A converts to B by cutting at 120 minutes, giving 189300=0.63\frac{189}{300} = 0.63; B cannot be converted back, since a yes hides the actual number of minutes. (d) With the same 300 students answering both, the gap of 216189=27216 - 189 = 27 students, or 9%, is a measurement discrepancy from the two wordings, not sampling variability, and it does not say which figure is right.

Problem 7

A class of 120 students reports a favorite lunch and a commute time. The lunch counts are pizza 48, salad 24, sandwich 36, and other 12, and the mean commute is 23.4 minutes. Four students each write one sentence about these data. Evaluate each sentence, correcting it where it is wrong.

(i) "The mean favorite lunch is pizza." (ii) "The counts 48, 24, 36, and 12 are numbers, so favorite lunch is a quantitative variable." (iii) "A histogram of favorite lunch shows that pizza is the most common." (iv) "The mean commute is 23.4 minutes, which is a decimal, so commute time is continuous."

Show the worked solution
  1. Check the totals before judging anything: 48+24+36+12=12048 + 24 + 36 + 12 = 120, so all 120 students are accounted for, and the proportions are 48120=0.40\frac{48}{120} = 0.40, 24120=0.20\frac{24}{120} = 0.20, 36120=0.30\frac{36}{120} = 0.30, and 12120=0.10\frac{12}{120} = 0.10, which add to 1.

  2. (i) Wrong word and wrong idea. Favorite lunch is categorical, so it has no mean at all: you cannot add pizza to salad. What the student means is the mode, the most common category, and pizza is indeed the mode at 48 students, or 40%. Corrected: "the most common favorite lunch is pizza, chosen by 48 of the 120 students."

  3. (ii) The reasoning confuses the tally with the thing tallied. The variable is favorite lunch, and its values are pizza, salad, sandwich, and other, which are labels. The numbers 48, 24, 36, and 12 are frequencies, and a frequency is always a number no matter what is being counted, so its being numeric proves nothing about the variable. The type belongs to what is counted; see count data. Favorite lunch stays categorical.

  4. (iii) Wrong graph. A histogram puts values into intervals on a number line, and its bars touch because the intervals touch. Pizza, salad, sandwich, and other are not intervals of anything and have no width, so the correct display is a bar graph with separated bars, or a pie chart of the four proportions.

  5. (iv) The conclusion is right and the reason is wrong, which on a free-response question earns nothing. Commute time is continuous because elapsed time fills an interval, not because a mean came out with a decimal in it.

  6. (iv) Counterexample, which is how to kill this reasoning outright: five students report 0, 1, 1, 2, and 3 siblings, so the mean is 0+1+1+2+35=75=1.4\frac{0 + 1 + 1 + 2 + 3}{5} = \frac{7}{5} = 1.4 siblings. Number of siblings is plainly discrete, and no student has 1.4 siblings. A mean does not have to be one of the possible values, so the shape of a mean says nothing about the type of the variable.

(i) Wrong: a categorical variable has no mean; pizza is the mode, at 48120=0.40\frac{48}{120} = 0.40. (ii) Wrong: 48, 24, 36, and 12 are frequencies, and the type belongs to what is counted, so favorite lunch is categorical. (iii) Wrong graph: use a bar graph or pie chart, since the four lunches are labels rather than intervals on a number line. (iv) Right conclusion, wrong reason: time is continuous because it fills an interval, and a decimal mean proves nothing, since 0, 1, 1, 2, 3 siblings averages 75=1.4\frac{7}{5} = 1.4 and siblings is discrete.

Problem 8

A city parks department studies its 5 public pools. On one Saturday, staff at every pool record, for each swimmer who enters: age in whole years, the size of the party the swimmer arrived with, membership status (member or day pass), which pool was entered, and the number of minutes the swimmer stayed. In all, 1,800 swimmers are recorded. The department reports a mean stay of 74.6 minutes, a mean party size of 2.3 people, and that 1,116 swimmers, or 62%, held day passes.

(a) Name the observational unit and give nn. (b) Classify all five variables. (c) Match each of the three reported numbers to its variable and name the kind of summary it is. (d) A staffer says "pool is quantitative, because we numbered the pools 1 through 5 and the mean pool number is 3.1." Evaluate that. (e) Is the 62% a parameter or a statistic?

Show the worked solution
  1. (a) One row is one swimmer entering on that Saturday, so the observational unit is a swimmer entry and n=1800n = 1800. It is not a pool: there are only 5 pools, and a pool is a value of one of the columns, not a row.

  2. (b) Age in whole years: quantitative and continuous. It is written as an integer because the form asks for whole years, but the quantity behind it is elapsed time, which fills an interval. Party size: quantitative and discrete, since parties come in 1, 2, 3, and onward with nothing between. Membership status: categorical, two labels. Pool entered: categorical, five labels. Minutes stayed: quantitative and continuous, since time fills an interval.

  3. (c) The mean stay of 74.6 minutes is a measure of center for the minutes column, a quantitative variable. The mean party size of 2.3 people is a measure of center for the party size column, also quantitative. The 62% is a proportion, which along with counts and the mode is the kind of summary a categorical variable allows, and it comes from the membership column: 11161800=0.62\frac{1116}{1800} = 0.62.

  4. (c) Note that 2.3 is not a possible party size, and that is fine. A mean of a discrete variable need not be one of its values, exactly as in problem 7.

  5. (d) The staffer has classified the labels, not the variable. Numbering the pools 1 through 5 assigns names written with digits, the same move as a jersey number in problem 2. A mean pool number of 3.1 identifies no pool and would change the moment the pools were renumbered, while not one swimmer's day would be different. Pool entered is categorical, and the summaries available are counts and proportions per pool, displayed with a bar graph.

  6. (e) Name the population before answering. If the population of interest is the swimmers who entered these 5 pools on that Saturday, then all 1,800 of them were measured, this is a census of that group, and 62% is a parameter. If the department means all swimmers across the season, those 1,800 are a sample from it, and 62% is a statistic.

  7. (e) Under the second reading, add the warning the design earns: the Saturday was picked for convenience rather than at random, and weekend swimmers plausibly buy day passes more often than weekday swimmers do, so 62% may sit above the season-long share. Which population you name decides both the word and whether the number can be trusted for it.

(a) The unit is one swimmer entry on that Saturday, n=1800n = 1800. (b) Age quantitative and continuous; party size quantitative and discrete; membership status categorical; pool entered categorical; minutes stayed quantitative and continuous. (c) 74.6 minutes is a mean of the minutes column, 2.3 people is a mean of the party size column, and 62% is a proportion from the membership column, 11161800=0.62\frac{1116}{1800} = 0.62. (d) Wrong: numbering the pools makes digits into labels, a mean pool number of 3.1 names no pool and changes under renumbering, so pool entered is categorical. (e) It depends on the population named: a parameter for the 1,800 swimmers of that Saturday, who were all measured, and a statistic if the target is all swimmers across the season, where one convenience-chosen Saturday may overstate the day pass share.