# The Investigative Question Revisited and Data Collection

> AP Statistics · AP Stats 2024-2026
> Source: https://www.owlsprep.com/study/ap-statistics-u10-the-investigative-question-revisited-and/

This module reviews framing testable investigative questions, selecting valid sampling and experimental data collection methods, and avoiding bias sources that cost points on AP Stats FRQs.

**Prerequisites:** [Basic definitions of population, sample, and parameter](https://www.owlsprep.com/study/ap-statistics-u3-population-sample-parameters-statistics/); [Introduction to observational studies vs experiments](https://www.owlsprep.com/study/ap-statistics-u3-observational-experiments-intro/)

## Learning objectives

- Distinguish between well-posed and ambiguous statistical investigative questions
- Match appropriate data collection methods to a given research question
- Identify sources of bias introduced by poor question framing or sampling design
- Justify data collection choices to earn full points on AP Statistics FRQ prompts

## Characteristics of a Well-Posed Investigative Question

A weak, vague investigative question makes all subsequent data collection and analysis useless, even if you perform calculations perfectly. AP exam graders explicitly award points for identifying and fixing poorly framed questions in FRQ responses.

**Well-Posed Statistical Investigative Question** — A question that explicitly names three components: 1) the exact target population, 2) the quantifiable outcome being measured, 3) the scope and time frame of the measurement.

**Worked example:** Rewrite the vague question 'Do teens exercise more?' to make it a well-posed investigative question.

1. First, identify the missing components: no defined population, no measurable exercise metric, no scope.
2. Add the population: '10th to 12th grade students at a US public high school'
3. Add the measurable outcome: 'average number of minutes of moderate to vigorous physical activity per week'
4. Add the scope: 'during the 2024-2025 academic school year'
5. Final well-posed question: 'What is the average number of minutes of moderate to vigorous physical activity per week for 10th to 12th grade students at a US public high school during the 2024-2025 academic year?'
6. Verify: this question can be answered with data, and two different researchers would collect comparable results.

**Check your understanding**

1. Which of the following is a well-posed investigative question?

   - Do dogs make better pets than cats?
   - What proportion of 8th graders in Ohio have at least one pet at home in 2025?
   - Are pets good for student mental health?
   - How many people like pets?

   *Why:* This option defines the population, measurable outcome, and scope clearly.

## Matching Data Collection Methods to Research Goals

No single data collection method works for every research goal. The AP exam will often ask you to select the most appropriate method for a given question, and justify your choice to earn full points.

**Comparing methods**

Below are the four core data collection methods you are expected to know for the AP exam, and their appropriate use cases:

- **Simple Random Sample Survey** — Randomly select individuals from a full sampling frame to answer survey questions
  - Pros: Minimizes selection bias, results generalize to the full population
  - Cons: Requires a complete, accurate list of all population members

- **Census** — Collect data from every single member of the target population
  - Pros: No sampling variability, produces exact population values
  - Cons: Extremely time-consuming and expensive for large populations

- **Completely Randomized Experiment** — Randomly assign subjects to treatment and control groups
  - Pros: Allows researchers to draw valid causal conclusions
  - Cons: Not always ethical or feasible for sensitive research questions

- **Observational Study** — Record existing behavior of subjects without applying treatments
  - Pros: Low effort, no ethical barriers for sensitive topics
  - Cons: Cannot prove causal relationships, only association

**Worked example:** A researcher wants to test if a new after-school math program improves student test scores. Which data collection method should they use, and why?

1. First, identify the research goal: test for a causal relationship between the program and test scores.
2. Eliminate unsuitable methods: a survey or observational study cannot prove causation, a census is unnecessary.
3. Select the correct method: a completely randomized experiment.
4. Justification: Randomly assign half of the volunteer students to take the new after-school program, and the other half to a standard after-school study hall. Compare final test scores between the two groups to isolate the effect of the program.

## Identifying Bias from Poor Question Framing

Even a well-designed sampling plan can be ruined by biased survey question wording. Leading questions, double-barreled questions, and ambiguous wording all introduce response bias that makes results untrustworthy.

> **warning**
>
> AP exam graders will deduct full points for identifying a leading question as 'fine' or 'unbiased'. This is one of the most commonly missed points on Unit 3 FRQs.

**Worked example:** Identify the bias in this survey question, and rewrite it to be neutral: 'Do you agree that the expensive new school lunch program that wastes taxpayer money is terrible?'

1. Identify the bias: this is a leading question that uses loaded negative language ('expensive', 'wastes taxpayer money', 'terrible') to push respondents to disagree with the lunch program.
2. Rewrite to remove loaded language: 'What is your opinion of the new school lunch program?'
3. Verify: the rewritten question does not guide respondents to any specific answer, and collects unbiased feedback.

**Exam command terms**

These are common command terms used for this topic on the AP exam, with their exact scoring expectations:

- **Identify** — Name the exact type of bias or flaw, no justification required *(Identify the bias in the survey question.)*

- **Explain** — Name the flaw AND describe how it affects the results *(Explain why the survey question produces biased results.)*

- **Justify** — Name your chosen data collection method AND link it explicitly to the research goal *(Justify your choice of data collection method for this study.)*

## Justifying Data Collection Choices for AP Exam Responses

Most students lose points on this topic not because they pick the wrong method, but because their justification is too vague. AP rubrics require you to make an explicit connection between your method and the research question's goals.

> **tip**
>
> A full-point justification always follows this structure: 1) Name your method, 2) State its key property, 3) Link that property to the specific goal of the question.

**Worked example:** A school administrator wants to estimate the proportion of all 2000 students at their school that support a new uniform policy. Write a full-point justification for using a stratified random sample by grade level.

1. 1) Name your method: I will use a stratified random sample, stratified by grade level.
2. 2) State its key property: Stratified random sampling ensures we select a representative number of students from each grade, rather than risking over-sampling one grade by chance.
3. 3) Link to the goal: This will produce a more accurate estimate of the proportion of all students that support the policy than a simple random sample, because opinions on uniforms often differ by grade level.

## Common pitfalls

- **Wrong:** Stating a vague question like 'do students like pizza?' without defining population or measurement
  - Why it fails: The question cannot produce replicable, quantifiable results, so you lose all 'validity' rubric points
  - Correct: Specify the exact population, measurable outcome, and scope: 'What proportion of 9th graders at this high school eat at least one slice of pizza per week?'
- **Wrong:** Choosing a convenience sample (e.g. surveying your friends) and claiming it is representative
  - Why it fails: Unaccounted for selection bias guarantees the sample cannot generalize to the target population
  - Correct: Explicitly name the bias, explain its effect, and propose a simple random sample or stratified random sample as an alternative
- **Wrong:** Confusing an observational study with an experiment to claim causal relationships
  - Why it fails: AP rubrics deduct points for incorrectly asserting causation without random assignment and control
  - Correct: State explicitly that observational studies can only show association, not causation
- **Wrong:** Writing leading survey questions that push respondents to a specific answer and not identifying the bias
  - Why it fails: Response bias from leading questions is a top 3 most commonly missed bias identification point on FRQs
  - Correct: Rewrite the question to be neutral: replace 'Do you agree the new lunch policy is unfair?' with 'What is your opinion of the new lunch policy?'
- **Wrong:** Forgetting to define the sampling frame when describing a data collection plan
  - Why it fails: Missing the sampling frame definition automatically loses 1 of 3 points on standard data design FRQs
  - Correct: Name the full list of population members you will draw your sample from (e.g. the school's full student enrollment roster) before describing random selection

## Cheatsheet

| Method | Best Use Case | Key Advantage | Key Limitation |
| --- | --- | --- | --- |
| Simple Random Sample | Generalizable population estimates | Minimizes selection bias | Requires complete sampling frame |
| Stratified Random Sample | Comparing subgroups of the population | Reduces sampling variability | Requires prior knowledge of subgroup membership |
| Completely Randomized Experiment | Testing causal relationships | Controls for confounding variables | Not always ethical or feasible |
| Voluntary Response Survey | No valid statistical use case | Low effort to distribute | Extreme nonresponse bias |
| Convenience Sample | Preliminary exploratory data only | Fast data collection | Cannot generalize to target population |

## What's next

Mastering how to frame strong investigative questions and select appropriate data collection methods is the foundational skill for all remaining AP Statistics units, as every inference, hypothesis test, and confidence interval you will learn later relies on data that was collected correctly. Poorly designed data collection makes all subsequent statistical calculations meaningless, even if you execute the math perfectly. This skill is tested in nearly every AP Stats exam's FRQ section, often as the first part of a multi-part question that moves from data design to inference. Next, you will build on this foundation by exploring sampling distributions, the core theoretical framework that underpins all inferential statistics on the exam.

---

From [OwlsPrep](https://www.owlsprep.com) — free study guides for A-Level, IB, AP and IGCSE, written against the official syllabus. Canonical page: https://www.owlsprep.com/study/ap-statistics-u10-the-investigative-question-revisited-and/
