How to describe a completely randomized design

By Jude Wallis · Published

Describe it in four moves: number the experimental units 1 to N, use a random number generator to choose which units get each treatment, state how many units land in each group, then name the response you measure and say you compare the groups on it.

AP Statistics: Unit 1 (topics 1.13 Experimental Design). The completely randomized design is named in topic 1.13 (Experimental Design), the last topic of Unit 1 in the Fall 2026 AP Statistics course, and describing one is a standard task on the free-response question about collecting data.

The four moves a complete description makes

A completely randomized design is described in four sentences, and they always do the same four jobs.

  1. Number the units. Label the NN experimental units 1 to NN.
  2. Assign at random. Use a random number generator to pick which numbered units go to each treatment, and say you ignore repeats.
  3. State the group sizes. Say exactly how many units end up in each treatment group.
  4. Measure and compare. Name the response variable you record on every unit, then say you compare the treatment groups on it.

Every one of the four is doing work. Numbering makes the units addressable. The generator makes the assignment something a stranger could carry out without asking you a question. The group sizes prove that all NN units got used. The response and the comparison are what makes the whole thing an experiment instead of a shuffling exercise.

The move students skip is the fourth. They build a careful randomization, then stop, and never say what gets measured or that the groups are compared. Write those two clauses even when they feel obvious.

What makes a design completely randomized

In a completely randomized design, the treatments are assigned to all of the experimental units completely at random. Nothing is grouped first. Every unit is in one big pool, and chance alone decides which treatment it receives.

That is the whole contrast with the other two designs in topic 1.13. A randomized block design sorts units into blocks that are alike on some extraneous variable first, then randomizes within each block. A matched pairs design is the two-treatment case of that. If your description sorts units into groups by a variable before randomizing, or pairs units up and randomizes inside each pair, you have described a blocked or matched pairs design, not a completely randomized one.

A well-designed experiment has four elements, and your write-up should make all four visible: comparison of at least two treatments, random assignment of treatments to units, replication (more than one unit per treatment), and direct control of other sources of variation, meaning you hold the rest of the conditions the same from unit to unit.

Random assignment has one job: to balance extraneous variables across the treatment groups on average, so that no systematic difference piles up in one group. Any single randomization can still come out uneven by chance, which is why blocking is worth the trouble when you already know a variable matters. That balancing on average is why an experiment can support a cause-and-effect conclusion at all, which is the point argued in experiments vs observational studies.

Write the randomization so someone else could run it

"Randomly assign the subjects to the two treatments" is not a description. It is a label for a description you did not write. The test to apply: could a stranger carry out your procedure without asking you a single question? If not, it is not finished.

Two mechanisms are enough for anything the exam asks.

  • Random number generator. Number the units 1 to NN. Generate random integers from 1 to NN, ignoring any number that repeats, until you have as many distinct numbers as the first treatment group needs. Those units get treatment 1. Keep going for treatment 2, and the units whose numbers never came up get the last treatment.
  • Physical randomization. Write each unit's number on an identical slip of paper, put the slips in a hat, mix thoroughly, and draw the numbers needed for each treatment group in turn.

Two clauses get dropped and both matter. Say that you ignore repeats, because a generator will hand you the same number twice and the procedure has to know what to do. Say where the leftover units go, because otherwise the last group never gets filled.

One mechanism to be careful with: flipping a coin for each subject. That is genuine random assignment, but it does not produce the group sizes you just promised. With 60 subjects, independent coin flips land on a 30 and 30 split only about 10% of the time. Use the generator or the hat when you have committed to fixed group sizes.

The diagram, and what has to be labeled

A diagram is a fast way to show the design, and boxes with arrows are fine. The boxes earn nothing on their own; the labels do. Here is what each stage has to say.

Stage of the diagramWhat the label must say
Unitshow many units there are, and that they are numbered
Random assignmentthe mechanism, not just the word "random"
Group 1its size and the treatment it receives
Group 2its size and the treatment it receives
Responsethe exact variable measured on every unit
Comparisonthe statistic compared across the groups

Filled in for a fertilizer trial with 40 garden plots and two fertilizers:

  • Units: 40 plots, numbered 1 to 40.
  • Random assignment: random number generator, integers 1 to 40, repeats ignored.
  • Group 1: 20 plots, new fertilizer.
  • Group 2: 20 plots, standard fertilizer.
  • Response: tomato yield in kilograms per plot at harvest.
  • Comparison: mean yield of the new-fertilizer plots against mean yield of the standard plots.

A diagram with two unlabeled boxes marked Group 1 and Group 2, an arrow, and no response variable describes nothing. If a question says describe, write the sentences; a diagram is a supplement to them, not a replacement.

The checklist to run before you move on

Read your paragraph back and tick these off. Each one is a separate thing a scorer is looking for.

  • The experimental units are named and numbered, and the total is stated.
  • A usable random mechanism appears, with repeats handled.
  • Every treatment is named, including the control or placebo if there is one.
  • Each group's size is stated, and the sizes add to the total.
  • The response variable is specific and measurable, with units where they exist.
  • The final sentence says the groups are compared on that response.
  • The conditions other than the treatment are held the same across groups.

This is mainly Question 1 territory on the free-response section, the multi-focus question on formulating questions and collecting data, though Question 4 draws on collecting data too. Questions are scored component by component, so a description that nails the randomization and forgets the response is not a near-miss, it is a description with a missing part. More on structuring free-response answers is in the AP Statistics FRQ guide.

Completely randomized or blocked: describe what was asked

Read the prompt for one signal: does it hand you a variable that is expected to affect the response and ask for a design that accounts for it? Words like the subjects differ a lot in baseline fitness, or the field has a wet end and a dry end, are an invitation to block. If nothing like that appears, describe the completely randomized design and stop.

Adding blocks nobody asked for does not earn extra credit. It lengthens the write-up, gives you more places to slip, and answers a question that was not on the page. The reverse mistake is worse: a question that names a nuisance variable and asks you to control for it wants blocking, and a completely randomized design does not do that job. The two designs and the sampling technique they get confused with are laid out in blocking vs stratifying.

One more distinction that decides sentences. Random assignment decides who gets which treatment, and it is what the words completely randomized refer to. Random selection decides who is in the study at all, and it is what lets you generalize to a population. A completely randomized design run on volunteers is still completely randomized; you just cannot generalize beyond people like those volunteers.

Mistakes that cost points

These are the ones that show up over and over.

  • Writing "randomly assign" with no mechanism. It is the single most common way to lose the randomization credit.
  • Forgetting the group sizes, which leaves the reader unable to tell whether all the units were used.
  • Never naming the response variable, or naming something vague like performance instead of score on the same 40-question test.
  • Leaving out the comparison sentence, so the design collects data but never answers anything.
  • Sorting units before randomizing (all the men in one group, the fittest half in another). That is not random assignment, and it reintroduces exactly the confounding the design exists to prevent.
  • Describing a block design when the question asked for a completely randomized one.
  • Confusing random assignment with random selection, and claiming the results generalize to everyone because the assignment was random.

Two treatments, 60 volunteers: the full write-up

A researcher wants to know whether caffeine shortens reaction time. She has 60 adult volunteers, a caffeinated drink, and an identical-tasting decaf drink. Describe a completely randomized design for this experiment.

  1. Number the units. The experimental units are the 60 volunteers. Label them 1 to 60.

  2. Set the group sizes. Two treatments split evenly gives 602=30\frac{60}{2} = 30 volunteers per group.

  3. Assign at random. Use a random number generator to produce random integers from 1 to 60, ignoring any repeat, until 30 distinct numbers have appeared. Those 30 volunteers drink the caffeinated drink. The remaining 6030=3060 - 30 = 30 volunteers drink the decaf drink.

  4. Control the rest. Give every volunteer the same drink volume, wait the same 30 minutes, and run the same computer reaction-time task in the same room. Neither the volunteers nor the person running the task knows which drink was served, which makes the study double-blind.

  5. Name the response. Record each volunteer's reaction time in milliseconds on that task.

  6. Compare. Compare the mean reaction time of the 30 caffeinated volunteers with the mean reaction time of the 30 decaf volunteers.

  7. Audit the four elements: comparison (two treatments), random assignment (the generator), replication (30 units per treatment, well above one), and direct control (same volume, same wait, same task, same room).

Number the 60 volunteers 1 to 60. Use a random number generator to generate random integers from 1 to 60, ignoring repeats, until 30 distinct numbers appear; those 30 volunteers receive the caffeinated drink and the other 30 receive the identical decaf drink. All 60 get the same volume, wait 30 minutes, and complete the same computer task, with neither the volunteers nor the person running the task knowing which drink was served. Record each volunteer's reaction time in milliseconds, then compare the mean reaction time of the caffeinated group with the mean reaction time of the decaf group.

Three treatments and a total that will not divide evenly

A teacher has 50 student volunteers and three review methods: flashcards, practice tests, and rereading notes. All students then take the same 40-question exam. Describe a completely randomized design, and handle the fact that 50 does not split into three equal groups.

  1. Number the units. Label the 50 students 1 to 50.

  2. Work out the group sizes. 50316.7\frac{50}{3} \approx 16.7, which is not a whole number, so make the groups as equal as possible: 17, 17, and 16. Check the total: 17+17+16=5017 + 17 + 16 = 50, so every student is used.

  3. Decide which method gets the group of 16, at random, so the choice is not made by preference. Number the three methods 1, 2, 3 and use a random number generator to pick one of them; that method gets the group of 16.

  4. Assign at random. Generate random integers from 1 to 50, ignoring repeats. The first 17 distinct numbers name the students assigned to one of the two methods that were not picked in step 3, the next 17 distinct numbers name the students assigned to the other method that was not picked, and the 16 students whose numbers never appeared get the method picked in step 3.

  5. Control the rest. Give every group the same amount of review time on the same material, and have all 50 students take the same 40-question exam under the same conditions.

  6. Name the response and compare. Record each student's score out of 40, then compare the mean score across the three method groups.

Label the students 1 to 50 and set group sizes of 17, 17, and 16, since 17+17+16=5017 + 17 + 16 = 50; number the three methods 1, 2, 3 and use a random number generator to pick which one gets the group of 16. Generate random integers from 1 to 50, ignoring repeats: the first 17 distinct numbers go to one of the two methods that were not picked, the next 17 to the other, and the 16 students never drawn take the method that was picked. Hold review time, material, and testing conditions the same for all groups, record each student's score on the same 40-question exam, and compare the mean scores of the three groups.

Frequently asked questions

Do the treatment groups have to be the same size?

No. A design is completely randomized as long as the treatments are assigned to all units at random. Equal sizes are the usual choice because they give the most precise comparison when the groups vary about the same amount. When the total does not divide evenly, make the groups as equal as possible, say so, and check that the sizes add up to the total.

Is writing "randomly assign the subjects to the treatments" enough?

No. That names the idea without describing it. Give a mechanism a stranger could carry out: number the units, generate random integers in that range, ignore repeats, and say which numbers go to which treatment and where the leftover units end up.

What is the difference between a completely randomized design and a randomized block design?

A completely randomized design assigns treatments to every unit from one pool, at random. A randomized block design first sorts units into blocks that are similar on an extraneous variable, then randomly assigns treatments inside each block, so every treatment appears in every block. Blocking buys precision when that variable really does affect the response.

Does a completely randomized design need a control group?

It needs at least two treatments to compare, and one of them may be a control group receiving a placebo or the current standard. Comparing two active treatments also counts. What you cannot do is run one group and call it an experiment, because there is nothing to compare it against.