Study Guide

Measures of central tendency

IB Mathematics AI SLΒ· 10 min read

1. Mean: Calculation for Raw and Grouped Dataβ˜…β˜…β˜†β˜†β˜†β± 10 min

πŸ“˜ Definition

Arithmetic Mean

xˉ\bar{x}

The sum of all data values divided by the number of values. For grouped data, it is an estimate calculated using midpoints of class intervals.

Example:

For data ,

πŸ“ Worked Example

The number of text messages 5 students sent in one hour are . Calculate the mean.

  1. 1

    Sum all values in the data set:

  2. 2
    12+15+8+10+15=6012 + 15 + 8 + 10 + 15 = 60
  3. 3

    Divide the sum by the number of observations :

  4. 4
    xˉ=605=12\bar{x} = \frac{60}{5} = 12
πŸ“ Worked Example

The table below shows heights of 30 students. Estimate the mean height.

  1. 1

    First, calculate midpoints and for each class:

  2. 2
    Class intervalMidpoint Frequency
    1454580
    155101550
    165111815
    1755875
  3. 3

    Sum frequencies and :

  4. 4
    βˆ‘f=30,βˆ‘fx=4820\sum f = 30, \quad \sum fx = 4820
  5. 5

    Calculate the estimated mean:

  6. 6
    xΛ‰=βˆ‘fxβˆ‘f=482030β‰ˆ160.7 cm\bar{x} = \frac{\sum fx}{\sum f} = \frac{4820}{30} \approx 160.7 \text{ cm}

Exam tip:

Always remember the mean from grouped data is only an estimate, not an exact value.

2. Median and Mode: Definition and Calculationβ˜…β˜…β˜†β˜†β˜†β± 8 min

πŸ“˜ Definition

Median

The middle value of an ordered data set. For observations: if is odd, it is the th value; if even, average of the th and th values.

πŸ“˜ Definition

Mode

The most frequently occurring value in a data set. For grouped data, the modal class is the class interval with the highest frequency.

πŸ“ Worked Example

For the ordered text message data , find the median and mode.

  1. 1

    We have (odd), so the median is the rd value. The 3rd value is 12, so median .

  2. 2

    The mode is the most frequent value. 15 occurs twice, so mode .

πŸ“ Worked Example

Add a 6th data point to the set above. Find the new median.

  1. 1

    Reorder the data: , (even).

  2. 2

    Average the 3rd and 4th values: . New median .

3. Selecting the Appropriate Measure of Central Tendencyβ˜…β˜…β˜…β˜†β˜†β± 10 min

The choice of measure depends on the type of data and shape of the distribution. Outliers and skewness affect which measure is most representative of the 'typical' value.

  • Mode: Only appropriate for nominal categorical data; can also be used for discrete numerical data.

  • Median: Appropriate for ordinal data, skewed numerical data, and data with outliers, as it is not affected by extreme values.

  • Mean: Appropriate for symmetric numerical data without outliers, it uses all data values but is sensitive to extreme values.

πŸ“ Worked Example

A business has 4 employees earning $2000 per week, and an owner earning $12000 per week. Which measure best describes the typical weekly wage? Justify.

  1. 1

    Ordered wages: . Mode = 2000, median = 2000, mean = .

  2. 2

    The owner's wage is an outlier that right-skews the distribution. The mean is pulled up to $4000, which is higher than 4 of 5 wages.

  3. 3

    Conclusion: The median (or mode) is best, as it is not affected by the outlier and reflects the typical wage.

4. Effects of Linear Data Transformationsβ˜…β˜…β˜…β˜†β˜†β± 7 min

When all values in a data set are transformed by adding or multiplying by a constant, all three measures of central tendency change in predictable ways:

  • Adding a constant to every value: mean, median and mode all increase by .

  • Multiplying every value by a constant : mean, median and mode are all multiplied by .

πŸ“ Worked Example

The mean height of a class is 161 cm, measured in centimeters. What is the new mean when converted to meters?

  1. 1

    Converting cm to meters requires multiplying every value by .

  2. 2

    All measures of central tendency are scaled by the same constant, so new mean = m.

5. Common Pitfalls

Wrong move:

Forgetting to order the data set before calculating the median

Why:

The median is only defined as the middle value of an ordered set; unordered data gives an incorrect result

Correct move:

Always sort your data from smallest to largest before finding the median

Wrong move:

Treating the estimated mean for grouped data as an exact value

Why:

We use midpoints to approximate the original raw data, which is not available for grouped sets

Correct move:

Always remember the mean from grouped data is an estimate, and round it appropriately

Wrong move:

Using the mean for skewed data with outliers

Why:

Extreme values pull the mean away from the center, leading to a misleading measure of typical value

Correct move:

Use the median for skewed data or data with extreme outliers

Wrong move:

Calculating a single mode for continuous raw data with all unique values

Why:

Continuous data rarely has repeated values, so a single mode is meaningless for raw continuous data

Correct move:

Use the median or mean for raw continuous data, or the modal class for grouped continuous data

Wrong move:

Expecting the median to always be a value from the original data set

Why:

For even-sized data sets, the median is the average of two middle values, which may not exist in the original data

Correct move:

When is even, average the two middle values even if the result is not in the original data

6. Quick Reference Cheatsheet

Measure

When to use

Formula/Rule

Mode

Nominal categorical data

Most frequent value / modal class

Median

Ordinal data, skewed data, data with outliers

Middle value of ordered data

Mean (raw)

Symmetric numerical data, no outliers

Mean (grouped)

Estimate for grouped continuous data

When this came up on past exams

AI-estimated based on syllabus patterns β€” cross-check with official past papers for accuracy. Use only as revision-focus signals.

  • 2025 Β· 1

    Calculate mean from grouped data

  • 2024 Β· 2

    Justify choice of central tendency measure

What's Next

Measures of central tendency are the foundation of descriptive statistics, and they are almost always paired with measures of spread to fully describe the shape and center of a data distribution. Understanding when to use each measure helps you avoid common mistakes in interpretation when analyzing real-world data, which is a core skill for IB AI SL statistics questions. You will often be asked to compare measures of central tendency between two or more data sets in exam questions, so building a solid understanding here will support all future statistical work from correlation to hypothesis testing.