Data representation: histograms, box plots, cumulative frequency
IB Mathematics: Applications and Interpretation SLΒ· Unit 4: Statistics and Probability, Topic 2Β· 15 min read
1. Histograms for Grouped Continuous Dataβ β ββββ± 5 min
Histogram
A graphical representation of grouped continuous data, where the area of each bar is proportional to the frequency of the class interval.
Example:
For a class 10-20 with frequency 15, frequency density =
Unlike bar charts for categorical data, histograms have no gaps between bars because the underlying data is continuous. When class widths are unequal, the height of each bar (frequency density) is not equal to frequency: only the area of the bar represents total frequency.
Calculate frequency density for the grouped mass data of 50 apples, and describe how to draw the histogram. Classes: 50-60 (f=8), 60-80 (f=22), 80-100 (f=20)
- 1
Calculate class width for each interval:
- 2
- 3
Calculate frequency density = frequency / class width for each interval:
- 4
- 5
Draw axes: mass on the horizontal axis, frequency density on the vertical axis. Draw each bar from the lower to upper class boundary, with height equal to the calculated frequency density.
Exam tip:
IB exam questions almost always use unequal class widths to test your understanding that area equals frequency, not height. Always calculate frequency density first for unequal classes.
2. Box Plots and Five-Number Summariesβ β ββββ± 5 min
Box Plot (Box-and-Whisker Plot)
A visual display of numerical data based on the five-number summary, used to compare distributions and identify skewness and outliers.
Example:
A symmetric distribution will have the median line in the centre of the box, and equal-length whiskers.
Box plots are the most common tool in IB exams for comparing two or more distributions. To get full marks, you must comment on both the centre (measured by the median) and the spread (measured by the interquartile range or range) of the distributions.
Find the five-number summary for the raw data set:
- 1
Confirm the data is sorted in ascending order (it is already sorted here, with values)
- 2
Find the median (Q2): position = term, so median =
- 3
Find Q1: median of the lower half of data (first 5 terms): term =
- 4
Find Q3: median of the upper half of data (last 5 terms): term =
- 5
Five-number summary: minimum = , Q1 = , median = , Q3 = , maximum =
Exam tip:
When comparing two box plots, always make two separate comments: one about average/centre (median) and one about variation/spread (IQR) to earn full marks.
3. Cumulative Frequency Diagramsβ β β βββ± 6 min
Cumulative Frequency Diagram
A plot of the running total of frequencies against upper class boundaries, used to estimate percentiles, quartiles and medians for grouped data.
To find any percentile, draw a horizontal line from the required cumulative frequency to the curve, then drop a vertical line down to the horizontal axis to read the estimated value. For total data points, median is at , Q1 at , Q3 at .
Estimate the median height for 100 students from the cumulative frequency data: Upper bound (cm): 150 (cf=12), 160 (cf=35), 170 (cf=68), 180 (cf=92), 190 (cf=100)
- 1
Median for 100 students is the 50th value, so we need the height at cumulative frequency = 50
- 2
Plot all (upper bound, cumulative frequency) points and join with a smooth curve
- 3
50 falls between cf=35 (160 cm) and cf=68 (170 cm). Interpolate to estimate:
- 4
4. Common Pitfalls
Wrong move:
Treating histogram bar height as frequency for unequal class widths
Why:
Confusion between histograms and bar charts leads to incorrect frequency calculations
Correct move:
Always calculate frequency density = frequency / class width for unequal classes, remember area = frequency
Wrong move:
Calculating quartiles for raw data without sorting first
Why:
Rushing through the question leads to incorrect positions and values for quartiles
Correct move:
Always sort raw data in ascending order before finding any percentiles or the five-number summary
Wrong move:
Plotting cumulative frequency against class midpoints
Why:
Confusion with histogram plotting conventions leads to incorrect percentile estimates
Correct move:
Always plot cumulative frequency against upper class boundaries, as it counts all values up to that boundary
Wrong move:
Only commenting on average when comparing box plots
Why:
Forgetting that comparison questions require comment on both centre and spread for full marks
Correct move:
Always make one comment about centre (median) and one comment about spread (IQR/range)
5. Quick Reference Cheatsheet
Graph Type | Key Features | Common Exam Uses | Quick Tip |
|---|---|---|---|
Histogram | Area = frequency, frequency density (y-axis), no gaps | Display grouped continuous data | Check for unequal class widths |
Box Plot | 5-number summary, shows outliers | Compare distributions, identify skewness | Compare median AND IQR |
Cumulative Frequency | Plot against upper class boundaries | Estimate quartiles/percentiles | Median = n/2, Q1 = n/4, Q3 = 3n/4 |
When this came up on past exams
AI-estimated based on syllabus patterns β cross-check with official past papers for accuracy. Use only as revision-focus signals.
- 2025 Β· Paper 1
Interpret box plot and compare distributions
- 2024 Β· Paper 2
Estimate quartiles from cumulative frequency
- 2023 Β· Paper 1
Construct histogram for unequal classes
Going deeper
What's Next
This sub-topic is the foundation of all descriptive statistics in IB AI SL, assessed in both Paper 1 and Paper 2. Mastering these data representations is essential for progressing to bivariate data analysis, correlation, and probability distributions. These skills also frequently appear in extended response questions that combine multiple statistics topics, so building a strong understanding early will help you with more complex topics later.
