The Investigative Question Revisited and Data Collection
AP StatisticsΒ· 12 min read
1. Characteristics of a Well-Posed Investigative Questionβ β ββββ± 3 min
A weak, vague investigative question makes all subsequent data collection and analysis useless, even if you perform calculations perfectly. AP exam graders explicitly award points for identifying and fixing poorly framed questions in FRQ responses.
Well-Posed Statistical Investigative Question
A question that explicitly names three components: 1) the exact target population, 2) the quantifiable outcome being measured, 3) the scope and time frame of the measurement.
Rewrite the vague question 'Do teens exercise more?' to make it a well-posed investigative question.
- 1
First, identify the missing components: no defined population, no measurable exercise metric, no scope.
- 2
Add the population: '10th to 12th grade students at a US public high school'
- 3
Add the measurable outcome: 'average number of minutes of moderate to vigorous physical activity per week'
- 4
Add the scope: 'during the 2024-2025 academic school year'
- 5
Final well-posed question: 'What is the average number of minutes of moderate to vigorous physical activity per week for 10th to 12th grade students at a US public high school during the 2024-2025 academic year?'
- 6
Verify: this question can be answered with data, and two different researchers would collect comparable results.
Which of the following is a well-posed investigative question?
Do dogs make better pets than cats?
What proportion of 8th graders in Ohio have at least one pet at home in 2025?
Are pets good for student mental health?
How many people like pets?
Reveal answer
What proportion of 8th graders in Ohio have at least one pet at home in 2025? βThis option defines the population, measurable outcome, and scope clearly.
2. Matching Data Collection Methods to Research Goalsβ β β βββ± 3 min
No single data collection method works for every research goal. The AP exam will often ask you to select the most appropriate method for a given question, and justify your choice to earn full points.
Below are the four core data collection methods you are expected to know for the AP exam, and their appropriate use cases:
Simple Random Sample Survey
Randomly select individuals from a full sampling frame to answer survey questions
+ Pros: Minimizes selection bias, results generalize to the full population
β Cons: Requires a complete, accurate list of all population members
Census
Collect data from every single member of the target population
+ Pros: No sampling variability, produces exact population values
β Cons: Extremely time-consuming and expensive for large populations
Completely Randomized Experiment
Randomly assign subjects to treatment and control groups
+ Pros: Allows researchers to draw valid causal conclusions
β Cons: Not always ethical or feasible for sensitive research questions
Observational Study
Record existing behavior of subjects without applying treatments
+ Pros: Low effort, no ethical barriers for sensitive topics
β Cons: Cannot prove causal relationships, only association
A researcher wants to test if a new after-school math program improves student test scores. Which data collection method should they use, and why?
- 1
First, identify the research goal: test for a causal relationship between the program and test scores.
- 2
Eliminate unsuitable methods: a survey or observational study cannot prove causation, a census is unnecessary.
- 3
Select the correct method: a completely randomized experiment.
- 4
Justification: Randomly assign half of the volunteer students to take the new after-school program, and the other half to a standard after-school study hall. Compare final test scores between the two groups to isolate the effect of the program.
3. Identifying Bias from Poor Question Framingβ β β βββ± 3 min
Even a well-designed sampling plan can be ruined by biased survey question wording. Leading questions, double-barreled questions, and ambiguous wording all introduce response bias that makes results untrustworthy.
Identify the bias in this survey question, and rewrite it to be neutral: 'Do you agree that the expensive new school lunch program that wastes taxpayer money is terrible?'
- 1
Identify the bias: this is a leading question that uses loaded negative language ('expensive', 'wastes taxpayer money', 'terrible') to push respondents to disagree with the lunch program.
- 2
Rewrite to remove loaded language: 'What is your opinion of the new school lunch program?'
- 3
Verify: the rewritten question does not guide respondents to any specific answer, and collects unbiased feedback.
4. Justifying Data Collection Choices for AP Exam Responsesβ β β β ββ± 3 min
Most students lose points on this topic not because they pick the wrong method, but because their justification is too vague. AP rubrics require you to make an explicit connection between your method and the research question's goals.
A school administrator wants to estimate the proportion of all 2000 students at their school that support a new uniform policy. Write a full-point justification for using a stratified random sample by grade level.
- 1
- Name your method: I will use a stratified random sample, stratified by grade level.
- 2
- State its key property: Stratified random sampling ensures we select a representative number of students from each grade, rather than risking over-sampling one grade by chance.
- 3
- Link to the goal: This will produce a more accurate estimate of the proportion of all students that support the policy than a simple random sample, because opinions on uniforms often differ by grade level.
5. Common Pitfalls
Wrong move:
Stating a vague question like 'do students like pizza?' without defining population or measurement
Why:
The question cannot produce replicable, quantifiable results, so you lose all 'validity' rubric points
Correct move:
Specify the exact population, measurable outcome, and scope: 'What proportion of 9th graders at this high school eat at least one slice of pizza per week?'
Wrong move:
Choosing a convenience sample (e.g. surveying your friends) and claiming it is representative
Why:
Unaccounted for selection bias guarantees the sample cannot generalize to the target population
Correct move:
Explicitly name the bias, explain its effect, and propose a simple random sample or stratified random sample as an alternative
Wrong move:
Confusing an observational study with an experiment to claim causal relationships
Why:
AP rubrics deduct points for incorrectly asserting causation without random assignment and control
Correct move:
State explicitly that observational studies can only show association, not causation
Wrong move:
Writing leading survey questions that push respondents to a specific answer and not identifying the bias
Why:
Response bias from leading questions is a top 3 most commonly missed bias identification point on FRQs
Correct move:
Rewrite the question to be neutral: replace 'Do you agree the new lunch policy is unfair?' with 'What is your opinion of the new lunch policy?'
Wrong move:
Forgetting to define the sampling frame when describing a data collection plan
Why:
Missing the sampling frame definition automatically loses 1 of 3 points on standard data design FRQs
Correct move:
Name the full list of population members you will draw your sample from (e.g. the school's full student enrollment roster) before describing random selection
6. Quick Reference Cheatsheet
Method | Best Use Case | Key Advantage | Key Limitation |
|---|---|---|---|
Simple Random Sample | Generalizable population estimates | Minimizes selection bias | Requires complete sampling frame |
Stratified Random Sample | Comparing subgroups of the population | Reduces sampling variability | Requires prior knowledge of subgroup membership |
Completely Randomized Experiment | Testing causal relationships | Controls for confounding variables | Not always ethical or feasible |
Voluntary Response Survey | No valid statistical use case | Low effort to distribute | Extreme nonresponse bias |
Convenience Sample | Preliminary exploratory data only | Fast data collection | Cannot generalize to target population |
When this came up on past exams
AI-estimated based on syllabus patterns β cross-check with official past papers for accuracy. Use only as revision-focus signals.
- 2023 Β· Set 1 FRQ
Evaluate survey question validity
- 2022 Β· Set 2 FRQ
Design data collection for study
- 2021 Β· Set 1 MCQ
Identify biased sampling method
What's Next
Mastering how to frame strong investigative questions and select appropriate data collection methods is the foundational skill for all remaining AP Statistics units, as every inference, hypothesis test, and confidence interval you will learn later relies on data that was collected correctly. Poorly designed data collection makes all subsequent statistical calculations meaningless, even if you execute the math perfectly. This skill is tested in nearly every AP Stats exam's FRQ section, often as the first part of a multi-part question that moves from data design to inference. Next, you will build on this foundation by exploring sampling distributions, the core theoretical framework that underpins all inferential statistics on the exam.
