Skills Beta

Mean, median, mode and range

5 min read · freeNot practiced

Why one measurement is never enough

Put a coin on a table and add water one drop at a time. The water piles up into a dome, held together by the hydrogen bonds between water molecules at the surface, until it finally spills. Count the drops and you might get 31. Dry the coin and do it again: 28. Again: 33. The coin did not change; small differences in drop size, how steady your hand was and where each drop landed changed the count.

This happens in every experiment. Repeated measurements of the same thing vary, so the honest result is not one number but a sample: the set of values you actually collected. The number of values in it is the sample size, written n. To report a sample you need two kinds of summary: where its values are centered, and how spread out they are.

Worked example: mean, median, mode and range

Data: drops held on a coin with plain water, measured eight times: 31, 28, 33, 30, 29, 35, 30, 32 drops.

Mean (add, then divide by n): x̄ = Σx ÷ n = (31 + 28 + 33 + 30 + 29 + 35 + 30 + 32) ÷ 8 = 248 ÷ 8 = 31.0 drops.

Median (sort, then take the middle): 28, 29, 30, 30, 31, 32, 33, 35. With 8 values there are two middle values, the 4th (30) and the 5th (31). Median = (30 + 31) ÷ 2 = 30.5 drops.

Mode (most frequent value): 30 appears twice, every other value once. Mode = 30 drops.

Range (largest − smallest): 35 − 28 = 7 drops.

On the exam's formula sheet the mean is written x̄ = (1/n) Σ xi: the symbol Σ ("sigma") means "add up all of them".

Mean or median? What an outlier does

A dot plot of eight drop counts: 13, 14, 14, 15, 15, 16, 16 and 31. The median, 15, sits in the middle of the cluster; the mean, 16.75, is pulled to the right toward the single high value of 31. The range runs from 13 to 31.
Figure 1. Eight drop counts with 2% soap. The median stays with the cluster; the mean is pulled toward the single high value. LevlPrep original diagram.

Now add soap to the water. Soap molecules crowd between water molecules at the surface, fewer hydrogen bonds hold the surface together, and the dome breaks sooner. One group got 14, 16, 13, 15, 31, 14, 16, 15 drops. Seven values sit between 13 and 16, and one is 31 (see the dot plot).

  • The mean is 134 ÷ 8 = 16.75 drops. Every value adds to the sum, so the 31 adds far more than the others and drags the mean above seven of the eight values.
  • The median is the average of the 4th and 5th sorted values: (15 + 15) ÷ 2 = 15 drops. It depends only on position, so the 31 could be 310 and the median would not move.

A value far from the rest is an outlier. When a data set has an outlier, or is lopsided with a long tail on one side, the median describes a typical value better than the mean. When the data are roughly symmetric, the two are close and the mean is usually reported because it uses all the information.

Comparing the summaries of center
SummaryHow to find itEffect of one outlierBest used when
MeanSum ÷ nPulled strongly toward itData roughly symmetric, no outliers
MedianMiddle of sorted dataBarely movesOutliers or lopsided data
ModeMost frequent valueNone (unless the outlier repeats)Counts or categories with a clear favorite

What to do with an outlier

An outlier is a question, not an answer. First check for a mistake: a copying error (31 written for 13), a dirty coin, a dropper that was swapped. If you find one, correct or remove the value and say so. If you find nothing, the value is real data. Keep it, report it, and give the median as well as the mean so a reader sees its effect. Deleting a value because it is inconvenient makes your results look more consistent than they were.

Spread: the range and why it matters

Two groups each report a mean of 25 drops. Group P got 24, 25, 25, 26, 25; group Q got 18, 31, 25, 22, 29. The centers match, but P's range is 26 − 24 = 2 drops and Q's is 31 − 18 = 13 drops. Q's measurements are far less consistent, so any single Q value could land well away from 25. A center without a spread hides this, so always report both.

The range is quick but uses only two values, so a single outlier inflates it (the 2% soap data have a range of 31 − 13 = 18 drops, mostly because of one value). Later skills replace it with measures of spread that use every value.

Types of data

Kinds of data you will meet
TypeWhat it isExamplesCan you average it?
Quantitative, discreteCounted in whole unitsDrops on a coin, number of seedsYes
Quantitative, continuousMeasured on a scaleMass (g), time (s), temperature (°C)Yes
Qualitative (categorical)Names or categoriesColor, spilled or notNo; count each category instead

A mean of counted drops can be a decimal (31.0, 16.75) even though each count is whole: the mean describes the sample, not one measurement.

Spot a mistake on this page?