Study Guide

Representation and Summary of Data (Edexcel IAL Maths S1)

MathematicsΒ· 2018 Specification Issue 3, S1 Β§2.1–§2.4Β· 25 min read

1. Interpreting Data Visualizations for Distribution Comparisonβ˜…β˜…β˜†β˜†β˜†β± 5 min

βœ“ Calculator OK

Exam questions almost always provide pre-drawn diagrams, so your focus is on extracting key statistics and comparing two or more distributions using context. Common diagrams include histograms, stem-and-leaf diagrams (including back-to-back variants) and box plots.

πŸ“˜ Definition

Box Plot

Diagram displaying minimum value, lower quartile (Q₁), median (Qβ‚‚), upper quartile (Q₃) and maximum value of a data set, clearly showing spread and central tendency.

πŸ“ Worked Example

The box plots below show test scores (out of 100) for Year 12 students studying Biology and Chemistry. Compare the distributions of scores for the two subjects.

  1. 1

    Compare central tendency: median Biology score = 68, median Chemistry score = 56, so typical Biology scores are 12 marks higher.

  2. 2

    Compare spread: IQR for Biology = 72 - 58 = 14, IQR for Chemistry = 62 - 43 = 19, so Biology scores are less varied.

  3. 3

    Compare range: total range for Biology = 88 - 32 = 56, total range for Chemistry = 92 - 22 = 70, so Chemistry scores have wider overall spread.

  4. 4

    Contextual conclusion: Overall, Biology students performed better on this test with more consistent scores than Chemistry students.

Exam tip:

Always reference context when comparing distributions, not just numerical values, to gain full marks.

2. Measures of Location & Codingβ˜…β˜…β˜…β˜†β˜†β± 7 min

βœ“ Calculator OK

Measures of location describe the central or typical value of a data set. You will calculate these for ungrouped and grouped discrete/continuous data, and use coding to simplify complex calculations.

πŸ“˜ Definition

Coding

y=xβˆ’aby = \frac{x - a}{b}

Linear transformation of raw data , where and are constants chosen to reduce the size of values you work with. The original mean is recovered using .

πŸ“ Worked Example

The height ( cm) of 30 students is coded using . The mean of the coded data is . Calculate the mean height of the students.

  1. 1

    Rearrange the coding formula to solve for in terms of :

  2. 2
    x=10y+150x = 10y + 150
  3. 3

    Apply the linear transformation to the mean of the coded data:

  4. 4
    xˉ=10yˉ+150=10(2.4)+150=24+150=174\bar{x} = 10\bar{y} + 150 = 10(2.4) + 150 = 24 + 150 = 174
  5. 5

    The mean height of the students is 174 cm.

3. Measures of Dispersionβ˜…β˜…β˜…β˜…β˜†β± 8 min

βœ“ Calculator OK

Measures of dispersion describe the spread of values in a data set. You will calculate range, interpercentile ranges, variance and standard deviation, and use simple interpolation for grouped data where percentiles fall between class boundaries.

πŸ“˜ Definition

Sample Standard Deviation

Square root of sample variance, measuring average deviation of values from the mean, in the same units as the original data. For coded data , .

πŸ“ Worked Example

For the same coded height data , the standard deviation of is 0.82. Calculate the standard deviation of the original height data, in cm.

  1. 1

    Recall that coding only scales standard deviation by the absolute value of the divisor : adding or subtracting a constant does not affect spread.

  2. 2
    sx=bΓ—sys_x = b \times s_y
  3. 3

    Substitute the given values (, ):

  4. 4
    sx=10Γ—0.82=8.2s_x = 10 \times 0.82 = 8.2
  5. 5

    The standard deviation of the original height data is 8.2 cm.

Exam tip:

When calculating interpercentile ranges for grouped data, always use class boundaries (not upper/lower class limits) for interpolation to avoid errors.

4. Skewness & Outliersβ˜…β˜…β˜…β˜†β˜†β± 5 min

βœ“ Calculator OK

Skewness describes the asymmetry of a distribution. You will be asked to interpret skewness from diagrams or summary statistics, and apply outlier rules explicitly given in the question (no default rule is required).

πŸ“˜ Definition

Positive (Right) Skew

Distribution with a long tail to the right, where mean > median > mode. Negative (left) skew has a long tail to the left, where mode > median > mean.

πŸ“ Worked Example

A distribution of monthly salary data has a mean of Β£3200, median of Β£2800 and mode of Β£2200. State the type of skewness shown, justifying your answer.

  1. 1

    Compare the values of mode, median and mean:

  2. 2
    2200<2800<3200β€…β€ŠβŸΉβ€…β€Šmode<median<mean2200 < 2800 < 3200 \implies \text{mode} < \text{median} < \text{mean}
  3. 3

    This ordering of measures corresponds to positive (right) skew, explained by a small number of very high salaries pulling the mean above the median.

5. Common Pitfalls

Wrong move:

Comparing distributions using only median values without mentioning spread or context

Why:

Exam questions require 2-3 comparison points for full marks, including central tendency, spread, and contextual interpretation

Correct move:

Always compare median, IQR/range, and add a 1-sentence contextual conclusion when comparing distributions

Wrong move:

Forgetting to reverse coding when calculating the original mean or standard deviation

Why:

Coded values are transformed, so you cannot use coded statistics directly as final answers for raw data questions

Correct move:

Use for mean and for standard deviation to recover original values after coding

Wrong move:

Using class limits instead of class boundaries when interpolating percentiles for grouped data

Why:

Class limits are the minimum and maximum values that can belong to a class, while boundaries are the points where classes meet, so using limits leads to incorrect interpolation results

Correct move:

Always convert grouped data classes to continuous boundaries before carrying out percentile interpolation

Wrong move:

Assuming a default outlier rule (e.g. 1.5Γ—IQR) when none is given

Why:

Edexcel IAL S1 explicitly states that all outlier rules are provided in the question, so using an unstated rule will give an incorrect answer and lose marks

Correct move:

Only apply the outlier rule exactly as written in the question text

Wrong move:

Confusing positive and negative skew by mixing up the order of mean, median and mode

Why:

Skew is named for the direction of the long tail, not the position of the peak, leading to easy mix-ups

Correct move:

Remember that positive skew has a tail to the right (higher values) so mean is pulled up above median, while negative skew has a tail to the left (lower values) so mean is pulled down below median

6. Quick Reference Cheatsheet

Concept

Formula/Rule

Exam Note

Mean (ungrouped)

For grouped data use midpoints as values

Coding Transformation ()

,

Adding/subtracting has no effect on standard deviation

IQR

Measures spread of middle 50% of data, unaffected by outliers

Standard Deviation

Same units as original data, variance is squared units

Skewness Interpretation

mode < median < mean = positive skew; mode > median > mean = negative skew

Skew is named for the direction of the long tail

Outliers

Apply rule given in question only

Do not use 1.5Γ—IQR rule unless explicitly stated

7. Frequently Asked

Do I need to memorize an outlier rule for S1 exams?

No: all outlier rules are explicitly provided in the question text for every Edexcel IAL S1 exam, so you do not need to learn any default rule such as 1.5Γ—IQR.

Will I be asked to draw histograms or box plots in the exam?

Direct drawing of diagrams is not a common focus; most questions ask you to interpret or compare given diagrams, or calculate statistics used to construct them.

Going deeper

What's Next

Now that you have mastered data representation and summary for Edexcel IAL S1, you are ready to move on to probability, the next core topic in the S1 specification. The concepts you have learned around distribution shape and spread will directly support your work on probability distributions, correlation and regression later in the course. Practice past paper questions on this topic to build speed and accuracy, paying close attention to command terms that ask you to compare, interpret or calculate values to maximize your marks.