Study Guide

Spearman's rank correlation coefficient

IB Mathematics Applications and Interpretation SLΒ· 30 min read

1. Core Definition and Purpose of Spearman's Rankβ˜…β˜…β˜†β˜†β˜†β± 8 min

[Block type renderer not yet implemented Β· content will appear once shipped]
πŸ“˜ Definition

Spearman's rank correlation coefficient

A statistical value between -1 and +1 that quantifies the strength and direction of monotonic association between two ranked bivariate variables

Example:

An value of 0.92 indicates a strong positive trend where higher values of one variable almost always correspond to higher values of the other

πŸ“ Worked Example

Identify which of the following datasets is better suited for Spearman's rank: A) Heights and weights of 20 teenagers, B) Customer satisfaction ratings (1-5) and product price for 15 items

  1. 1

    Evaluate dataset A: Heights and weights typically have a roughly linear relationship, so Pearson's r is appropriate

  2. 2

    Evaluate dataset B: Satisfaction ratings are ordinal, and the relationship to price is not guaranteed to be linear, so Spearman's rank is the correct choice

2. Step-by-Step Calculation for Untied Ranksβ˜…β˜…β˜…β˜†β˜†β± 12 min

βœ“ Calculator OK

rs=1βˆ’6βˆ‘di2n(n2βˆ’1)r_s = 1 - \frac{6\sum d_i^2}{n(n^2 - 1)}
  1. Rank all x values from 1 (smallest) to n (largest)

  2. Rank all corresponding y values from 1 to n

  3. Calculate the difference between the x rank and y rank for every paired observation

  4. Square each value to get

  5. Sum all values to get

  6. Substitute the sum into the standard Spearman's rank formula to get

πŸ“ Worked Example

Calculate for the following 4 paired data points: (2,5), (4,7), (6,10), (8,9)

  1. 1

    Rank x values: 2=1, 4=2, 6=3, 8=4

  2. 2

    Rank y values: 5=1,7=2,9=3,10=4

  3. 3

    Calculate values: 1-1=0, 2-2=0, 3-4=-1, 4-3=1

  4. 4

    Calculate values: 0, 0, 1, 1. Sum = 2

  5. 5

    Substitute into formula:

βœ“ Quick check
  1. What is the maximum possible value of for n=5?

    Reveal answer
    40 β€”

    Maximum sum occurs when ranks are perfectly reversed, sum of squares is 16+9+4+1+0 = 30? No wait, for n=5 reversed ranks, d_i values are 4,2,0,-2,-4, sum of squares 16+4+0+4+16=40, correct.

3. Adjusting for Tied Ranks in Datasetsβ˜…β˜…β˜…β˜…β˜†β± 10 min

[Block type renderer not yet implemented Β· content will appear once shipped]
πŸ“ Worked Example

Assign ranks to the following x values: [12, 15, 15, 18]

  1. 1

    The smallest value 12 occupies position 1, so rank = 1

  2. 2

    The two 15 values occupy positions 2 and 3. Average rank = (2+3)/2 = 2.5 for both entries

  3. 3

    The largest value 18 occupies position 4, so rank = 4

4. Interpretation and Comparison to Pearson's rβ˜…β˜…β˜…β˜†β˜†β± 7 min

Methods compared

Spearman's Rank $r_s$

Uses ranked data, works for non-linear trends, suitable for ordinal and skewed data

+ Pros: Robust to outliers, no normality assumption

βˆ’ Cons: Does not measure linear relationship strength

Pearson's $r$

Uses raw continuous data, only measures linear association

+ Pros: More precise for linear datasets

βˆ’ Cons: Sensitive to outliers, requires normal distribution

5. Common Pitfalls

Wrong move:

Using Pearson's r formula instead of Spearman's for non-linear monotonic data

Why:

Pearson's r will underestimate the true strength of association for non-linear trends

Correct move:

Use Spearman's rank whenever you only expect a consistent increasing/decreasing trend, not a straight line relationship

Wrong move:

Assigning different ranks to tied data points

Why:

Produces artificially large values and incorrect final

Correct move:

Assign the mean of all occupied rank positions to every tied entry

Wrong move:

Rounding intermediate values to 1 decimal place early

Why:

Cumulative rounding error makes your final fall outside the acceptable mark range

Correct move:

Keep all intermediate values unrounded until you calculate the final

Wrong move:

Claiming proves a perfect linear relationship

Why:

Spearman's r_s only measures perfect monotonic association, not linearity

Correct move:

State that means one variable strictly increases as the other increases, regardless of the shape of the trend

Wrong move:

Applying Spearman's rank to fewer than 5 paired observations

Why:

The coefficient becomes extremely sensitive to random variation and is not statistically meaningful

Correct move:

Confirm you have at least 6 paired data points before performing the calculation

6. Quick Reference Cheatsheet

Step

Untied Ranks Action

Tied Ranks Action

Interpretation Rule

1

Rank x values 1 to n

Assign average rank to tied x values

= perfect positive monotonic association

2

Rank y values 1 to n

Assign average rank to tied y values

= no monotonic association

3

Calculate

Use adjusted ranks to find

= perfect negative monotonic association

4

Sum all values

Sum adjusted values

0.7 < ≀ 1 indicates strong correlation

7. Frequently Asked

Can I use my GDC to calculate Spearman's rank directly?

Most modern GDC models can compute r_s automatically, but you must show full working for the rank and difference steps in Paper 2 to earn method marks, even if your final value matches the correct answer.

When this came up on past exams

AI-estimated based on syllabus patterns β€” cross-check with official past papers for accuracy. Use only as revision-focus signals.

  • 2023 Β· Paper 2

    Rank correlation for travel time data

  • 2022 Β· Paper 1

    Interpret given Spearman's r_s value

  • 2021 Β· Paper 2

    Calculate r_s for student score data

Going deeper

  • official_guideIB Math AI SL Formula SheetReference the Spearman's rank formula section

What's Next

Now that you have mastered Spearman's rank correlation, you can apply this non-parametric skill to a wide range of exam scenarios, from ordinal survey data to skewed environmental measurements. This topic is frequently tested in both Paper 1 and Paper 2, so ensure you practice full calculation sequences without skipping steps to avoid losing method marks. Next, you will build on your understanding of correlation to learn linear regression, which lets you create predictive models for bivariate datasets with confirmed linear association. You will also explore hypothesis testing to confirm if an observed correlation value is statistically significant rather than due to random chance.