AP Statistics Quiz: Introducing Statistics Do Those Points Align
19 questions · exam conditions
0:00
Introducing Statistics Do Those Points AlignQuestion 1 of 19

After collecting data on the number of years a person has been running and their time to complete a 10k race, a coach finds a negative linear association in her sample. Which question is most directly related to the statistical concept of inference for regression slope?

What is the fastest 10k time recorded in the sample data?
How strong is the evidence from this sample that a negative linear relationship exists for all runners of this type?
How many runners were included in the sample to ensure the results are valid?
Is the relationship between years running and 10k time stronger than the relationship between age and 10k time?
← Back to quizzes

AP Statistics Quiz

AP Statistics Quiz: Introducing Statistics Do Those Points Align

Practice Introducing Statistics Do Those Points Align in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Introducing Statistics Do Those Points Align, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

After collecting data on the number of years a person has been running and their time to complete a 10k race, a coach finds a negative linear association in her sample. Which question is most directly related to the statistical concept of inference for regression slope?

  1. What is the fastest 10k time recorded in the sample data?
  2. How strong is the evidence from this sample that a negative linear relationship exists for all runners of this type? (correct answer)
  3. How many runners were included in the sample to ensure the results are valid?
  4. Is the relationship between years running and 10k time stronger than the relationship between age and 10k time?

Explanation: Inference for the slope is about assessing the strength of the evidence provided by the sample to make a conclusion about the population. This question directly asks about the strength of evidence for a linear relationship at the population level.

Question 2

Imagine a population where for any given value of an explanatory variable xx, there is a distribution of corresponding values for a response variable yy. If a linear relationship exists, the means of these distributions fall on a straight line. Why would individual points from this population not fall perfectly on that line?

  1. Because any sample drawn from the population will have sampling error.
  2. Because the relationship between xx and yy is actually curved, not linear.
  3. Because the slope of the population regression line is not equal to zero.
  4. Because of inherent random variation in the response variable yy that is unrelated to the explanatory variable xx. (correct answer)

Explanation: This describes the fundamental model for linear regression. The line captures the mean response, but individual responses vary randomly around that mean. This variation is the error term (ϵ\epsilon) in the model y=α+βx+ϵy = \alpha + \beta x + \epsilon.

Question 3

A consumer advocacy group studied the relationship between the sugar content (in grams) and the consumer rating (on a scale of 1 to 100) for a random sample of breakfast cereals. They found a sample slope of -1.5. A primary question for statistical inference would be to determine if this negative slope is:

  1. the correct slope for this sample, since the least-squares method provides a unique line for the data.
  2. exactly equal to the population slope, which can be confirmed by increasing the sample size.
  3. due to sampling variability, or if it provides evidence of a true negative linear relationship in the population of all cereals. (correct answer)
  4. a result of a non-linear relationship that should be corrected by transforming the data before analysis.

Explanation: Statistical inference aims to distinguish between a result that could have happened by chance (sampling variability) and a result that is statistically significant, suggesting a real effect or relationship in the larger population.

Question 4

An economist models the relationship between a country's GDP and its average life expectancy using a linear regression model on data from a sample of countries. After fitting the line, she observes that the residuals are all positive for low and high GDPs and negative for mid-range GDPs. What does this pattern suggest about the linear model?

  1. A linear model may not be appropriate because the variation of points around the line is non-random. (correct answer)
  2. The relationship is linear, but the sample size was too small to accurately estimate the true slope.
  3. The association between GDP and life expectancy is negative, which contradicts the model's assumptions.
  4. The variation of points around the regression line is purely random, which is expected in a good model.

Explanation: A distinct pattern in the residual plot, such as a curve, indicates that the relationship between the variables is likely non-linear. The variation of points around the line is therefore not random, and a simple linear model is probably not the best fit for the data.

Question 5

A trainer at a gym wants to investigate if there is a positive linear relationship between the number of hours a person spends at the gym per week and the amount of weight they can lift. Data is collected from a random sample of gym members. From an inferential standpoint, what is the fundamental question the trainer is trying to answer?

  1. Does the observed positive slope from the sample provide convincing evidence that the true slope for all gym members is also positive? (correct answer)
  2. Is the slope calculated from the sample data greater than the intercept calculated from the sample data?
  3. Can the number of hours spent at the gym be used to perfectly predict the amount of weight a person can lift for this sample?
  4. What is the specific increase in lifting ability for each additional hour spent at the gym for the members in the sample?

Explanation: This question properly frames the inferential task: using the evidence from a sample (the observed positive slope) to make a claim about the larger population (whether the true slope is positive).

Question 6

A researcher calculates a sample regression line for the relationship between caffeine intake and hours of sleep for a random sample of students. She knows that if she took another random sample, she would likely get a different sample slope. How does this concept of sampling variability in the slope affect her conclusions about the true relationship?

  1. It means her first sample is likely biased, and she should take many more samples to find the correct slope.
  2. It introduces uncertainty, which is why she would use a confidence interval to estimate the true slope. (correct answer)
  3. It proves that no true linear relationship exists between caffeine intake and hours of sleep in the population.
  4. It suggests that the relationship is non-linear and a different model is required to capture the true pattern.

Explanation: The fact that sample slopes vary from sample to sample (sampling variability) means there is uncertainty in any single sample's slope as an estimate of the population slope. A confidence interval is a statistical tool designed to account for this uncertainty by providing a range of plausible values for the true population slope.

Question 7

A biologist believes there is a true linear relationship between the height of a certain species of plant and the amount of a specific nutrient in the soil. She collects data from a random sample of 50 plants. What does the least-squares regression line calculated from this sample represent?

  1. An estimate of the true linear relationship for all plants of this species. (correct answer)
  2. The exact linear relationship for all plants of this species.
  3. The only possible linear model that could be used to describe the data for this specific sample.
  4. Proof that a linear relationship exists between the two variables for all plants of this species.

Explanation: A least-squares regression line created from sample data is a statistic that serves as an estimate of the unknown population regression line. It represents our best guess for the true linear relationship based on the available sample data.

Question 8

A researcher investigates the relationship between hours of weekly exercise and resting heart rate for adults at a large company. They take a random sample of 30 adults and calculate the slope of the least-squares regression line. If they were to take a second, independent random sample of 30 adults from the same company, which of the following is most likely to be true?

  1. The slope of the regression line for the second sample would be different from the first sample's slope due to sampling variability. (correct answer)
  2. The slope of the regression line for the second sample would be exactly the same as the first sample's slope because the underlying population is the same.
  3. The slope of the regression line for the second sample would be exactly equal to the true population slope.
  4. The intercept of the regression line for the second sample would be the same as the first, but the slope would be different.

Explanation: Due to sampling variability, different random samples from the same population will almost certainly produce different values for sample statistics, including the slope and intercept of the least-squares regression line.

Question 9

A sociologist repeatedly takes random samples of size 40 from a large population of workers to study the relationship between years of education and annual income. For each sample, the slope of the least-squares regression line is calculated. Which of the following best describes the resulting collection of all the calculated sample slopes?

  1. A population distribution of income.
  2. A distribution of residuals.
  3. A sampling distribution of the slope. (correct answer)
  4. A scatterplot of income versus education.

Explanation: A sampling distribution is the distribution of a statistic (in this case, the slope) from all possible samples of a given size. The collection of slopes from repeated samples forms the sampling distribution of the slope, which shows how the sample slope varies.

Question 10

The AP Statistics course description mentions "variation in points' positions relative to a theoretical line." In the context of inference for the slope of a regression line, what is this "theoretical line"?

  1. The sample regression line, which is calculated from the observed data.
  2. A horizontal line with a slope of zero, representing the null hypothesis.
  3. The population regression line, which represents the true mean relationship between the variables. (correct answer)
  4. A line with a slope of 1 and an intercept of 0, representing a perfect one-to-one relationship.

Explanation: The "theoretical line" is the true, but usually unknown, population regression line (μy=α+βx\mu_y = \alpha + \beta x) that we try to estimate with our sample data. The variation of points around this line is the population error.

Question 11

An analyst studies the relationship between a stock's daily price change and the daily change in a major market index for a random sample of 60 trading days. The calculated slope of the least-squares regression line is b=0.02b = 0.02. What is the main purpose of conducting a hypothesis test for the slope in this context?

  1. To check if the conditions for linear regression were met for the sample of 60 trading days.
  2. To prove that the true relationship between the stock and the market index is exactly β=0.02\beta = 0.02.
  3. To determine if the small sample slope of 0.02 is statistically different from 0, or if it could have occurred by chance. (correct answer)
  4. To calculate the exact price change of the stock for a given change in the market index for days not in the sample.

Explanation: A hypothesis test for the slope is used to determine if the observed sample slope provides statistically significant evidence of a linear relationship in the population. The test assesses whether the sample slope is far enough from the hypothesized value (usually 0) that it's unlikely to be due to random chance.

Question 12

A student is looking at a scatterplot of fuel efficiency (miles per gallon) versus vehicle weight (pounds) for a random sample of 25 car models. The points form a general downward linear trend. Which question addresses the concept of variation that is central to beginning an inference procedure for the slope?

  1. Could the observed negative association have occurred by chance if there is actually no linear relationship between fuel efficiency and weight for all cars? (correct answer)
  2. What is the exact fuel efficiency for a car that weighs 3,000 pounds based on the sample data?
  3. Is the range of the vehicle weights in the sample large enough to accurately represent all car models?
  4. How much does the predicted fuel efficiency change for each additional pound of vehicle weight in this sample?

Explanation: This question gets at the heart of inference for slope: determining whether an observed pattern in a sample is strong enough to conclude it represents a real pattern in the population, or if it could just be a result of random sampling variation.

Question 13

A professor fits a linear model to data from a random sample of students, relating hours spent studying to exam scores. The scatterplot of the data shows points scattered randomly above and below the calculated least-squares regression line. What is the most appropriate interpretation of this observation?

  1. The model is flawed because it fails to predict the exact exam score for each student.
  2. The scatter indicates that there is no relationship between studying and exam scores because the points do not fall perfectly on a line.
  3. The random scatter supports the appropriateness of the linear model, with deviations representing natural variability. (correct answer)
  4. The data must have been collected with measurement error, which caused the points to deviate from the line.

Explanation: A key condition for linear regression is that the residuals (deviations from the line) are randomly scattered. This random scatter doesn't invalidate the model; instead, it confirms that a linear model is appropriate and that the leftover variation is simply random error.

Question 14

Two different researchers are studying the same population to determine the linear relationship between age and blood pressure. Researcher A takes a random sample of 50 individuals, and Researcher B takes a different random sample of 50 individuals. Both calculate the least-squares regression line. Which statement is most likely to be true?

  1. The two sample slopes will be identical because they are estimating the same population slope.
  2. One of the researchers must have made a mistake, as random samples should produce the same results.
  3. The sample slope from Researcher A will be different from Researcher B's, but both are estimates of the same population slope. (correct answer)
  4. The average of the two sample slopes is guaranteed to be the true population slope.

Explanation: This illustrates the concept of sampling variability. Because each researcher drew a different random sample, their calculated sample statistics (like the slope) will likely differ. Both slopes, however, serve as estimates for the single, true population slope.

Question 15

A scientist examines a scatterplot from a single random sample of data. The points appear to align closely to a straight line with a positive slope. What can the scientist conclude from this plot alone before performing any formal inference?

  1. There is a definite positive linear relationship in the population.
  2. There appears to be a positive linear association in the sample, suggesting there might be one in the population. (correct answer)
  3. The slope of the population regression line is positive and significantly different from zero.
  4. A different sample would produce the exact same scatterplot and regression line.

Explanation: A scatterplot from a sample can only describe the sample. We can observe an association and describe it, but we cannot make a definitive conclusion about the population without performing formal statistical inference. The sample suggests a possible relationship in the population.

Question 16

In many studies involving inference for the slope of a regression line, the null hypothesis is that the true slope is zero (β=0\beta = 0). What is the practical implication if one fails to reject this null hypothesis?

  1. It proves that there is no relationship of any kind between the two variables in the population.
  2. The two variables have a perfect correlation of zero in the sample data.
  3. There is not enough evidence to conclude that a linear relationship exists between the two variables in the population. (correct answer)
  4. The sample data must have been collected improperly, leading to a biased result.

Explanation: Failing to reject the null hypothesis means the data from the sample are not strong enough to rule out the possibility that the true slope is zero. We don't prove the null is true, but rather conclude there is insufficient evidence for the alternative hypothesis (that a linear relationship exists).

Question 17

In the context of simple linear regression, what does the population regression line, μy=α+βx\mu_y = \alpha + \beta x, describe?

  1. The mean value of the response variable yy for all individuals in the population that have a specific value of the explanatory variable xx. (correct answer)
  2. The exact value of the response variable yy for every individual in the population for a given value of the explanatory variable xx.
  3. The relationship between the explanatory and response variables for a particular sample drawn from the population.
  4. The predicted value of the explanatory variable xx based on a specific value of the response variable yy.

Explanation: The population regression line describes the average response (μy\mu_y) for a given explanatory value (xx). Individual responses will vary around this mean value.

Question 18

When examining a scatterplot of bivariate quantitative data from a population, what does the random scatter of points around the true population regression line represent?

  1. Natural variation in the response variable that is not explained by the linear relationship with the explanatory variable. (correct answer)
  2. A mistake in the calculation of the regression line, as all points should fall perfectly on the line in the population.
  3. The sampling error associated with estimating the population regression line from a single sample.
  4. Evidence that the relationship between the variables is not truly linear and a different model should be considered.

Explanation: In a population, individual data points will not all fall on the regression line. The vertical deviation of each point from the line represents the random error or natural variability in the response variable that cannot be accounted for by the linear model.

Question 19

The entire premise of conducting statistical inference for the slope of a least-squares regression line is based on what fundamental concept regarding the data collection process?

  1. The relationship between the two variables must be perfectly linear, with no deviation of points from the line.
  2. The data come from a random sample or a randomized experiment. (correct answer)
  3. The sample size must be greater than 10 percent of the population size to ensure accuracy.
  4. The explanatory variable must be the direct cause of the changes observed in the response variable.

Explanation: All methods of statistical inference rely on the principle of randomness in data collection (either random sampling or random assignment in an experiment). This randomness allows us to use probability to describe the behavior of sample statistics and make conclusions about the population.