AP Statistics · Unit 1 · 20-30% of the exam · ~26 class periods

Exploring One-Variable Data and Collecting Data

Unit 1 explores one-variable data (graphs, shape, center, spread, outliers) and how to collect data through sampling and experiments. At 20-30% of the multiple-choice section and about 26 class periods, it carries the heaviest weight of the five units.

AP Statistics: Unit 1 (topics 1.1 Introducing Statistics: What Can We Learn from Data?, 1.2 Variables, 1.3 Tabular Representation and Summary Statistics for One Categorical Variable, 1.4 Graphical Representations for One Categorical Variable, 1.5 Graphical Representations for One Quantitative Variable, 1.6 Descriptions for One Quantitative Variable Distributions, 1.7 Summary Statistics for One Quantitative Variable, 1.8 Graphical Representations of Summary Statistics for One Quantitative Variable, 1.9 Comparisons of the Distributions for One Quantitative Variable, 1.10 The Investigative Question Revisited and Data Collection, 1.11 Random Sampling, 1.12 Potential Problems with Sampling, 1.13 Experimental Design). Unit 1 is the first of the five units in the redesigned AP Statistics course (effective Fall 2026, first exam May 2027), and none of the topics removed in the redesign came from it, so its content is intact.

What Unit 1 of AP Statistics covers

Unit 1, Exploring One-Variable Data and Collecting Data, is the first of the five units in the redesigned AP Statistics course (effective Fall 2026, first exam May 2027). It carries the heaviest exam weight of the five units at 20-30% of the multiple-choice section, and the College Board plans about 26 class periods for it.

The unit has two halves. The first is about exploring data you already have: building graphs for one categorical or one quantitative variable, then describing and comparing distributions. The second is about collecting data well: writing a clear investigative question, choosing a random sampling method, spotting bias, and designing an experiment. Both halves return throughout the rest of the course.

Everything rests on four words. A population is all the individuals of interest (size NN), a sample is the subset you measure (size nn), a parameter summarizes a population, and a statistic summarizes a sample. Keeping these straight separates a full-credit answer from a vague one.

The 13 topics in Unit 1, one by one

  • 1.1 Introducing Statistics: What Can We Learn from Data? A statistical study collects data from a sample to answer an investigative question about a larger population, and you learn the core vocabulary along with how to pose a question that stays fixed once the data arrive.
  • 1.2 Variables. An observational unit is what you measure; a variable is a trait that changes across units. You sort variables into categorical or quantitative (and quantitative into discrete or continuous), and tell a parameter from a statistic.
  • 1.3 Tabular Representation and Summary Statistics for One Categorical Variable. Frequency tables show counts per category and relative frequency tables show proportions. Percentages, proportions, and ratios all carry the same information.
  • 1.4 Graphical Representations for One Categorical Variable. Bar charts and pie charts display the counts or proportions for one categorical variable, and you use them to compare two or more data sets on the same variable.
  • 1.5 Graphical Representations for One Quantitative Variable. Histograms, stem-and-leaf plots, and dotplots show a quantitative variable's distribution while keeping values in order, and choices such as histogram bin width change the picture.
  • 1.6 Descriptions for One Quantitative Variable Distributions. You describe a quantitative distribution by shape (skewed left, skewed right, or symmetric), center, and spread, plus outliers, gaps, and clusters, and you label peaks as unimodal, bimodal, or uniform.
  • 1.7 Summary Statistics for One Quantitative Variable. You calculate center (mean, median), position (quartiles, percentiles), and spread (range, IQR, standard deviation), and explain why the median and IQR resist outliers while the mean and standard deviation do not.
  • 1.8 Graphical Representations of Summary Statistics for One Quantitative Variable. The five-number summary becomes a boxplot whose box holds the middle 50% of the data, and the mean sitting above the median usually signals right skew (below it, left skew).
  • 1.9 Comparisons of the Distributions for One Quantitative Variable. You compare two or more distributions of the same variable with graphs and numerical summaries, and you use z-scores to compare relative positions within or between distributions.
  • 1.10 The Investigative Question Revisited and Data Collection. You split an investigative question into parts that guide collection, analysis, and conclusions, and you identify a census, experiment, observational study, survey, and confounding variables.
  • 1.11 Random Sampling. You identify sampling methods (simple random, stratified, cluster, systematic) and sampling with or without replacement, and justify which method fits a question.
  • 1.12 Potential Problems with Sampling. Bias is a systematic error that pushes a statistic consistently off the true parameter. You identify voluntary response, undercoverage, nonresponse, and response bias.
  • 1.13 Experimental Design. A well-designed experiment compares at least two treatments, assigns them at random, replicates, and controls other variation. You identify completely randomized, randomized block, and matched pairs designs and know that random assignment supports cause-and-effect conclusions.

Describing distributions with shape, center, and spread

When you describe the distribution of one quantitative variable, the CED asks for shape, center, and variability (spread), plus any unusual features such as outliers, gaps, or clusters, always in context. Many students recall this with the acronym SOCS: Shape, Outliers and gaps, Center, Spread.

  • Shape: skewed right, skewed left, or roughly symmetric, and unimodal (one peak), bimodal (two peaks), or approximately uniform.
  • Center: the mean, written xˉ\bar{x} (the sample mean), or the median (the middle value). See mean vs median for when to use each.
  • Spread: the range, the interquartile range (IQR), or the standard deviation, written ss for a sample. Work the formula in standard deviation by hand.
  • Unusual features: the AP outlier rule flags any value more than 1.5×IQR1.5 \times \text{IQR} above Q3Q_3 (the third quartile) or below Q1Q_1 (the first quartile); see the 1.5 IQR rule.

Two habits protect points: name the variable and its units every time, and never stop at center. The most common lost point on these questions is describing only the middle of the data and skipping shape, spread, or an obvious outlier. Practice in the descriptive statistics sandbox.

Sampling methods and experimental design

The second half of Unit 1 is about how data are collected, because the method decides what conclusions you may draw. Two ideas do the heavy lifting, and the exam expects you to keep them apart.

Random selection means using chance to pick who is in the sample, which lets you generalize your results to the population. The four named methods are the simple random sample (every possible sample of size nn is equally likely), the stratified random sample (sample within similar groups), the cluster random sample (take every individual in a few randomly chosen groups), and the systematic random sample (a random start, then every kth individual). Bias is a systematic error that pushes a statistic consistently off the parameter, in forms such as voluntary response, undercoverage, nonresponse, and response bias.

Random assignment is a different idea that belongs to experiments, where a researcher assigns treatments to experimental units. A well-designed experiment compares at least two treatments, assigns them at random, uses replication (more than one unit per treatment), and controls other sources of variation. Random assignment is what lets you conclude a treatment caused a difference in the response. Common designs are the completely randomized, randomized block, and matched pairs designs; for imposing treatments versus only observing, read experiments vs observational studies.

On the exam, use random selection when you generalize to a population and random assignment when you claim cause and effect.

How Unit 1 is tested on the AP exam

The exam is three hours long. Section I is 42 multiple-choice questions worth 50% of the score in 90 minutes, and Section II is 4 free-response questions worth 50% in 90 minutes, with each free-response question worth 10 points, or 12.5% of the exam score. A graphing calculator with statistical capabilities is expected on both sections, and a formula sheet and tables are provided.

Unit 1 is 20-30% of the multiple-choice section on its own. Its data-collection skills also fall under the Collect Data practice, which is 20-30% of the multiple-choice section. On the free-response side, Question 1 focuses on formulating questions and collecting data, and Question 4 draws on collecting data, analyzing data, and interpreting results, so sampling and experimental-design reasoning can appear directly in the written questions. Describing and comparing distributions shows up both in multiple-choice items and as parts of longer free-response problems.

For a walkthrough of the written section, see the AP Statistics FRQ guide.

Study priorities for Unit 1

If you are short on time, put your effort here:

  • Describe distributions completely and in context. Cover shape, center, spread, and unusual features every time.
  • Learn the 1.5 IQR outlier rule and the five-number summary, and be able to turn them into a boxplot.
  • Know when the mean or the median is the better measure of center, and why the median and IQR resist outliers while the mean, range, and standard deviation do not.
  • Work through a standard deviation by hand at least once so the formula is not a mystery.
  • Learn the four random sampling methods, the four kinds of bias, and the elements of a well-designed experiment.
  • Practice the experiments vs observational studies distinction and the random selection versus random assignment wording.

From here the course moves to Unit 2 on probability and random variables. You can see all five units on the AP Statistics hub.

Describe a distribution and check for outliers

Nine students report their commute time to school in minutes: 8, 10, 12, 15, 15, 18, 20, 22, 45. Find the five-number summary, use the 1.5 IQR rule to check for outliers, and compare the mean and median to describe the shape.

  1. Put the data in order (already ordered): 8, 10, 12, 15, 15, 18, 20, 22, 45. There are n = 9 values.

  2. The median is the middle (5th) value: median = 15 minutes.

  3. Lower half, the four values below the median: 8, 10, 12, 15. The first quartile is the median of these: Q1=(10+12)/2=11Q_1 = (10 + 12) / 2 = 11 minutes.

  4. Upper half, the four values above the median: 18, 20, 22, 45. The third quartile is Q3=(20+22)/2=21Q_3 = (20 + 22) / 2 = 21 minutes.

  5. Five-number summary: minimum 8, Q1Q_1 11, median 15, Q3Q_3 21, maximum 45.

  6. Interquartile range: IQR=Q3Q1=2111=10\text{IQR} = Q_3 - Q_1 = 21 - 11 = 10.

  7. Outlier fences: 1.5×IQR=1.5×10=151.5 \times \text{IQR} = 1.5 \times 10 = 15. Upper fence =Q3+15=21+15=36= Q_3 + 15 = 21 + 15 = 36. Lower fence =Q115=1115=4= Q_1 - 15 = 11 - 15 = -4.

  8. Compare each value to the fences: 45 > 36, so 45 is a high outlier. No value is below -4, so there are no low outliers.

  9. Mean: add the values, 8+10+12+15+15+18+20+22+45=1658 + 10 + 12 + 15 + 15 + 18 + 20 + 22 + 45 = 165. The sample mean is xˉ=165/9=18.33\bar{x} = 165 / 9 = 18.33 minutes (rounded to two decimals).

  10. The mean (18.33) is larger than the median (15), and the single large value 45 stretches the right tail.

The five-number summary is 8, 11, 15, 21, 45 minutes. The value 45 is a high outlier by the 1.5 IQR rule. Because the mean (18.33 minutes) exceeds the median (15 minutes) and the long tail is on the high side, the distribution is skewed right.

Compare two values with z-scores

Priya scored 85 on a biology test where the class mean was 78 and the standard deviation was 4. On a chemistry test, Maya scored 90 where the class mean was 84 and the standard deviation was 8. Treating each class as the full population, who did better relative to their own class?

  1. A z-score measures how many standard deviations a value is above or below the mean: z=xiμσz = \frac{x_i - \mu}{\sigma}, where xix_i is the data value, μ\mu (mu) is the population mean, and σ\sigma (sigma) is the population standard deviation.

  2. Priya: z=(8578)/4=7/4=1.75z = (85 - 78) / 4 = 7 / 4 = 1.75.

  3. Maya: z=(9084)/8=6/8=0.75z = (90 - 84) / 8 = 6 / 8 = 0.75.

  4. Priya's score sits 1.75 standard deviations above her class mean; Maya's sits 0.75 above hers.

Priya has the higher z-score (1.75 versus 0.75), so she performed better relative to her own class, even though Maya's raw score of 90 was higher than Priya's 85.

Frequently asked questions

How much of the AP Statistics exam is Unit 1?

Unit 1 is 20-30% of the multiple-choice section, the heaviest weight of the five units in the Fall 2026 course. The College Board allots about 26 class periods to it. Its data-collection skills also appear in the free-response section.

What does SOCS stand for in AP Statistics?

SOCS is a memory aid for describing the distribution of a quantitative variable: Shape, Outliers and gaps, Center, Spread. The CED expects all of these, plus the variable named in context. Describing only the center is a common way to lose points.

What is the difference between random selection and random assignment?

Random selection means using chance to choose who is in your sample, which lets you generalize results to the population. Random assignment means using chance to assign treatments in an experiment, which lets you draw cause-and-effect conclusions. The exam expects you to use the correct term for each situation.

Do I need a calculator for Unit 1?

Yes. A graphing calculator with statistical capabilities is expected on both sections of the exam, and a formula sheet and tables are provided. For Unit 1 you will mostly use it for summary statistics like the mean, standard deviation, and five-number summary.

Every topic in Unit 1

  1. 1.1Introducing Statistics: What Can We Learn from Data?
  2. 1.2Variables
  3. 1.3Tabular Representation and Summary Statistics for One Categorical Variable
  4. 1.4Graphical Representations for One Categorical Variable
  5. 1.5Graphical Representations for One Quantitative Variable
  6. 1.6Descriptions for One Quantitative Variable Distributions
  7. 1.7Summary Statistics for One Quantitative Variable
  8. 1.8Graphical Representations of Summary Statistics for One Quantitative Variable
  9. 1.9Comparisons of the Distributions for One Quantitative Variable
  10. 1.10The Investigative Question Revisited and Data Collection
  11. 1.11Random Sampling
  12. 1.12Potential Problems with Sampling
  13. 1.13Experimental Design