Study Guide

The Investigative Question Revisited and Data Collection

AP StatisticsΒ· 12 min read

1. Characteristics of a Well-Posed Investigative Questionβ˜…β˜…β˜†β˜†β˜†β± 3 min

A weak, vague investigative question makes all subsequent data collection and analysis useless, even if you perform calculations perfectly. AP exam graders explicitly award points for identifying and fixing poorly framed questions in FRQ responses.

πŸ“˜ Definition

Well-Posed Statistical Investigative Question

A question that explicitly names three components: 1) the exact target population, 2) the quantifiable outcome being measured, 3) the scope and time frame of the measurement.

πŸ“ Worked Example

Rewrite the vague question 'Do teens exercise more?' to make it a well-posed investigative question.

  1. 1

    First, identify the missing components: no defined population, no measurable exercise metric, no scope.

  2. 2

    Add the population: '10th to 12th grade students at a US public high school'

  3. 3

    Add the measurable outcome: 'average number of minutes of moderate to vigorous physical activity per week'

  4. 4

    Add the scope: 'during the 2024-2025 academic school year'

  5. 5

    Final well-posed question: 'What is the average number of minutes of moderate to vigorous physical activity per week for 10th to 12th grade students at a US public high school during the 2024-2025 academic year?'

  6. 6

    Verify: this question can be answered with data, and two different researchers would collect comparable results.

βœ“ Quick check
  1. Which of the following is a well-posed investigative question?

    • Do dogs make better pets than cats?

    • What proportion of 8th graders in Ohio have at least one pet at home in 2025?

    • Are pets good for student mental health?

    • How many people like pets?

    Reveal answer
    What proportion of 8th graders in Ohio have at least one pet at home in 2025? β€”

    This option defines the population, measurable outcome, and scope clearly.

2. Matching Data Collection Methods to Research Goalsβ˜…β˜…β˜…β˜†β˜†β± 3 min

No single data collection method works for every research goal. The AP exam will often ask you to select the most appropriate method for a given question, and justify your choice to earn full points.

Methods compared

Below are the four core data collection methods you are expected to know for the AP exam, and their appropriate use cases:

Simple Random Sample Survey

Randomly select individuals from a full sampling frame to answer survey questions

+ Pros: Minimizes selection bias, results generalize to the full population

βˆ’ Cons: Requires a complete, accurate list of all population members

Census

Collect data from every single member of the target population

+ Pros: No sampling variability, produces exact population values

βˆ’ Cons: Extremely time-consuming and expensive for large populations

Completely Randomized Experiment

Randomly assign subjects to treatment and control groups

+ Pros: Allows researchers to draw valid causal conclusions

βˆ’ Cons: Not always ethical or feasible for sensitive research questions

Observational Study

Record existing behavior of subjects without applying treatments

+ Pros: Low effort, no ethical barriers for sensitive topics

βˆ’ Cons: Cannot prove causal relationships, only association

πŸ“ Worked Example

A researcher wants to test if a new after-school math program improves student test scores. Which data collection method should they use, and why?

  1. 1

    First, identify the research goal: test for a causal relationship between the program and test scores.

  2. 2

    Eliminate unsuitable methods: a survey or observational study cannot prove causation, a census is unnecessary.

  3. 3

    Select the correct method: a completely randomized experiment.

  4. 4

    Justification: Randomly assign half of the volunteer students to take the new after-school program, and the other half to a standard after-school study hall. Compare final test scores between the two groups to isolate the effect of the program.

3. Identifying Bias from Poor Question Framingβ˜…β˜…β˜…β˜†β˜†β± 3 min

Even a well-designed sampling plan can be ruined by biased survey question wording. Leading questions, double-barreled questions, and ambiguous wording all introduce response bias that makes results untrustworthy.

πŸ“ Worked Example

Identify the bias in this survey question, and rewrite it to be neutral: 'Do you agree that the expensive new school lunch program that wastes taxpayer money is terrible?'

  1. 1

    Identify the bias: this is a leading question that uses loaded negative language ('expensive', 'wastes taxpayer money', 'terrible') to push respondents to disagree with the lunch program.

  2. 2

    Rewrite to remove loaded language: 'What is your opinion of the new school lunch program?'

  3. 3

    Verify: the rewritten question does not guide respondents to any specific answer, and collects unbiased feedback.

4. Justifying Data Collection Choices for AP Exam Responsesβ˜…β˜…β˜…β˜…β˜†β± 3 min

Most students lose points on this topic not because they pick the wrong method, but because their justification is too vague. AP rubrics require you to make an explicit connection between your method and the research question's goals.

πŸ“ Worked Example

A school administrator wants to estimate the proportion of all 2000 students at their school that support a new uniform policy. Write a full-point justification for using a stratified random sample by grade level.

  1. 1
    1. Name your method: I will use a stratified random sample, stratified by grade level.
  2. 2
    1. State its key property: Stratified random sampling ensures we select a representative number of students from each grade, rather than risking over-sampling one grade by chance.
  3. 3
    1. Link to the goal: This will produce a more accurate estimate of the proportion of all students that support the policy than a simple random sample, because opinions on uniforms often differ by grade level.

5. Common Pitfalls

Wrong move:

Stating a vague question like 'do students like pizza?' without defining population or measurement

Why:

The question cannot produce replicable, quantifiable results, so you lose all 'validity' rubric points

Correct move:

Specify the exact population, measurable outcome, and scope: 'What proportion of 9th graders at this high school eat at least one slice of pizza per week?'

Wrong move:

Choosing a convenience sample (e.g. surveying your friends) and claiming it is representative

Why:

Unaccounted for selection bias guarantees the sample cannot generalize to the target population

Correct move:

Explicitly name the bias, explain its effect, and propose a simple random sample or stratified random sample as an alternative

Wrong move:

Confusing an observational study with an experiment to claim causal relationships

Why:

AP rubrics deduct points for incorrectly asserting causation without random assignment and control

Correct move:

State explicitly that observational studies can only show association, not causation

Wrong move:

Writing leading survey questions that push respondents to a specific answer and not identifying the bias

Why:

Response bias from leading questions is a top 3 most commonly missed bias identification point on FRQs

Correct move:

Rewrite the question to be neutral: replace 'Do you agree the new lunch policy is unfair?' with 'What is your opinion of the new lunch policy?'

Wrong move:

Forgetting to define the sampling frame when describing a data collection plan

Why:

Missing the sampling frame definition automatically loses 1 of 3 points on standard data design FRQs

Correct move:

Name the full list of population members you will draw your sample from (e.g. the school's full student enrollment roster) before describing random selection

6. Quick Reference Cheatsheet

Method

Best Use Case

Key Advantage

Key Limitation

Simple Random Sample

Generalizable population estimates

Minimizes selection bias

Requires complete sampling frame

Stratified Random Sample

Comparing subgroups of the population

Reduces sampling variability

Requires prior knowledge of subgroup membership

Completely Randomized Experiment

Testing causal relationships

Controls for confounding variables

Not always ethical or feasible

Voluntary Response Survey

No valid statistical use case

Low effort to distribute

Extreme nonresponse bias

Convenience Sample

Preliminary exploratory data only

Fast data collection

Cannot generalize to target population

When this came up on past exams

AI-estimated based on syllabus patterns β€” cross-check with official past papers for accuracy. Use only as revision-focus signals.

  • 2023 Β· Set 1 FRQ

    Evaluate survey question validity

  • 2022 Β· Set 2 FRQ

    Design data collection for study

  • 2021 Β· Set 1 MCQ

    Identify biased sampling method

What's Next

Mastering how to frame strong investigative questions and select appropriate data collection methods is the foundational skill for all remaining AP Statistics units, as every inference, hypothesis test, and confidence interval you will learn later relies on data that was collected correctly. Poorly designed data collection makes all subsequent statistical calculations meaningless, even if you execute the math perfectly. This skill is tested in nearly every AP Stats exam's FRQ section, often as the first part of a multi-part question that moves from data design to inference. Next, you will build on this foundation by exploring sampling distributions, the core theoretical framework that underpins all inferential statistics on the exam.