Data types, sampling, bias
IB Mathematics: Applications and Interpretation HLΒ· 30 min read
1. Classifying Data Typesβ βββββ± 10 min
Data Classification
All statistical data is classified by type, which determines what analysis can be performed. Data is first split into qualitative (categorical) and quantitative (numerical). Quantitative data is further split into discrete (countable) and continuous (measurable).
Example:
Eye color = qualitative, number of cars owned = discrete, distance traveled = continuous
Classify each of the following as qualitative, discrete quantitative, or continuous quantitative: (a) The weight of oranges in a shipment, (b) The number of correct answers on a 20-question quiz, (c) The brand of phone owned by a student, (d) The temperature of a cup of coffee.
- 1
Recall the definitions: Qualitative = categorical non-numerical, discrete = distinct countable values, continuous = any value in an interval.
- 2
(a) Weight is a numerical measurement that can take any value between a range (e.g. 120.5g, 120.54g).
- 3
Classification: (a) = continuous quantitative.
- 4
(b) Number of correct answers is numerical, only whole number values between 0 and 20.
- 5
Classification: (b) = discrete quantitative.
- 6
(c) Phone brand is a non-numerical category.
- 7
Classification: (c) = qualitative.
- 8
(d) Temperature can take any value within a range, so it is a continuous measurement.
- 9
Classification: (d) = continuous quantitative.
Test your understanding
What type of data is the number of goals scored in a soccer season?
Qualitative
Discrete quantitative
Continuous quantitative
Reveal answer
Discrete quantitative βGoals are counted as whole numbers, so they are discrete, even though they are numerical.
2. Common Sampling Methodsβ β ββββ± 15 min
Population vs Sample
A population is the full set of individuals/items you want to draw conclusions about in a study. A sample is a smaller subset selected to collect data from, when studying the full population is impractical.
Sampling Method | Description | Typical Use Case |
|---|---|---|
Simple Random | Every member of the population has equal chance of selection | Small, accessible populations |
Systematic | Select every th member from an ordered list of the population | Large populations with a pre-existing list |
Stratified | Divide population into strata (groups by a key characteristic), sample from each stratum | Ensure proportional representation of important subgroups |
Cluster | Divide population into geographically similar clusters, randomly select entire clusters to sample | Large, geographically dispersed populations |
Convenience | Select easily accessible members | Pilot studies or preliminary research |
A researcher wants to study student satisfaction with campus food, and needs to ensure representation of undergraduate, graduate, and international student groups. What sampling method is most appropriate, and how would they implement it?
- 1
The study requires representation of predefined subgroups, so stratified random sampling is the most appropriate method.
- 2
Step 1: Split the full population of students into 3 strata: undergraduates, graduates, international students.
- 3
Step 2: Calculate the proportion of the total population each stratum makes up, to get proportional sample sizes for each group.
- 4
Step 3: Use simple random sampling to select the required number of students from each stratum, then send the satisfaction survey.
3. Identifying Sampling Biasβ β ββββ± 15 min
Sampling Bias
Bias is a systematic error in sampling that makes the sample unrepresentative of the target population. Common types include selection bias, voluntary response bias, non-response bias, and response bias.
A researcher conducts a survey about weekly exercise habits by standing outside a gym and asking people entering to complete the survey. What type of bias is present, and how will it affect results?
- 1
The sample is selected only from people entering a gym, which systematically excludes people who do not go to this gym.
- 2
This is selection bias: people who go to a gym exercise more regularly on average than the general population.
- 3
The results will systematically overestimate the average amount of weekly exercise for the general population.
4. Reliability and Validity of Data (AHL)β β β βββ± 15 min
Beyond avoiding bias, an HL study must produce data that is both reliable (consistent when repeated) and valid (actually measures what it claims to). These are distinct: a measurement can be highly reliable yet invalid β for example a miscalibrated scale that always reads 2 kg too heavy gives consistent but wrong values.
Reliability vs Validity
Reliability is about consistency of repeated measurements; validity is about whether the method measures the intended characteristic. A method must be reliable to be valid, but reliability alone does not guarantee validity.
Example:
A stopped clock is perfectly reliable (always shows the same time) but not valid (it does not measure the actual time).
Test | Type | How it works |
|---|---|---|
Test-retest | Reliability | Give the same test to the same group on two occasions and compare; strong agreement indicates high reliability. |
Parallel forms | Reliability | Give two equivalent versions of the test to the same group and compare; consistent scores indicate reliability without memory effects. |
Content validity | Validity | Expert judgement that the items fully cover the topic or construct being measured. |
Criterion-related validity | Validity | Compare results against an established external benchmark (criterion), e.g. correlating a new aptitude test with later job performance. |
A school designs a new 20-question questionnaire to measure student wellbeing. Describe one reliability test and one validity test the school could use.
- 1
Reliability (test-retest): Give the same questionnaire to the same group of students two weeks apart, then check whether each student's scores are close on the two occasions. Strong agreement indicates the questionnaire is reliable.
- 2
Validity (content validity): Ask wellbeing experts to review the 20 questions and confirm they cover all relevant aspects of wellbeing (emotional, social, physical) rather than, say, only measuring academic stress.
A chi-squared goodness-of-fit test checks whether the number of calls per minute follows a Poisson distribution. After combining tail classes there are 6 classes, and the Poisson mean was estimated from the sample. Find the degrees of freedom.
- 1
Start with (number of classes) minus 1:
- 2
- 3
One parameter (the Poisson mean) was estimated from the data, so subtract a further 1:
- 4
5. Common Pitfalls
Wrong move:
Confusing stratified and cluster sampling
Why:
Both methods split the population into groups, so they are often mixed up
Correct move:
Remember: Stratified sampling groups by shared characteristic, you sample from every group. Cluster sampling groups by location, you sample entire groups at random.
Wrong move:
Calling shoe size or IQ continuous data
Why:
They are numerical measurements, so they are often misclassified as continuous
Correct move:
Any variable that only takes distinct, separated values (whole/half shoe sizes, whole IQ points) is discrete quantitative.
Wrong move:
Assuming a large sample eliminates bias
Why:
Bias is caused by the sampling method, not sample size
Correct move:
A large biased sample is still systematically unrepresentative. Always check the sampling method for bias first, regardless of sample size.
Wrong move:
Calling coded categorical data quantitative
Why:
Coded qualitative data is stored as a number, leading to misclassification
Correct move:
Classify data based on what it measures, not how it is stored. A number that just labels a category (e.g. 1 = New York, 2 = London) is still qualitative.
6. Quick Reference Cheatsheet
Concept | Quick Reference |
|---|---|
Data Types | Qualitative = categorical; Discrete = countable values; Continuous = any measurable value |
Sampling Methods | Simple Random = equal chance; Systematic = every nth; Stratified = sample each subgroup; Cluster = sample whole groups; Convenience = easy access |
Common Bias Types | Selection = groups excluded; Voluntary = self-selected respondents; Non-response = missing participants; Response = inaccurate answers |
What's Next
This foundational sub-topic underpins all statistical work in IB AI HL. Correct classification of data and identification of bias is required for every topic from descriptive statistics to hypothesis testing, and exam questions regularly test these foundational concepts in both paper 1 and paper 2.
After mastering data types, sampling, and bias, you will move on to organizing and visualizing data, then learn to summarize data with measures of center and spread, before moving on to probability and inferential statistics.
