What this quiz covers
This quiz focuses on Linear Regression Models, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
A streaming service sampled 20 users and recorded age (x, years, from 13 to 62) and average hours streamed per week (y). A scatterplot with the least-squares regression line is shown, with fitted equation y^=18.0−0.15x. The purpose of the linear model is to describe the linear association and predict typical weekly streaming time from age within the observed range. Which interpretation of the model is correct?
AP Statistics Quiz
Practice Linear Regression Models in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.
This quiz focuses on Linear Regression Models, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.
A streaming service sampled 20 users and recorded age (x, years, from 13 to 62) and average hours streamed per week (y). A scatterplot with the least-squares regression line is shown, with fitted equation y^=18.0−0.15x. The purpose of the linear model is to describe the linear association and predict typical weekly streaming time from age within the observed range. Which interpretation of the model is correct?
Explanation: This question tests interpretation of age-related regression models. The equation y^=18.0−0.15x predicts weekly streaming hours from user age. The correct interpretation (A) states that for each additional year of age, predicted weekly streaming time decreases by about 0.15 hours on average. This properly acknowledges the associative and average nature of the relationship. Choice B incorrectly implies causation from aging itself. Choice C attempts to apply the model to newborns (age 0), far outside the observed range of 13-62 years. Choice D extrapolates to 120 years, well beyond the data range. Choice E completely misinterprets what the intercept represents. Linear models should only be used within the range of observed data - extrapolation to extreme ages produces unreliable and often nonsensical predictions.
A botanist measured the amount of fertilizer applied to a plot (x, in grams) and the plant height after 6 weeks (y, in centimeters) for 9 plots, with fertilizer amounts ranging from 0 to 40 grams. A least-squares regression line is y^=12.4+0.48x. The purpose of this linear model is to summarize the linear association and predict typical plant height for fertilizer amounts within the observed range. Which interpretation of the model is correct?
Explanation: The skill involves interpreting y^=12.4+0.48x for plant height from fertilizer (x from 0 to 40 grams). The slope shows 0.48 cm taller per gram on average in the range. Choice A correctly interprets without causation. Distractor B assumes causation for every plant. Choice C treats the intercept as exact for zero fertilizer. Limitation: no extrapolation beyond data, as choice D does to 200 grams. Slope doesn't equal explained variation.
An environmental scientist measured water temperature (x, in °C, from 6 to 24) and dissolved oxygen (y, mg/L) at 12 sites in a river. The least-squares regression line predicting dissolved oxygen from temperature is y^=12.1−0.18x. The purpose of this linear model is to describe the linear association and predict typical dissolved oxygen for temperatures in the observed range. Which interpretation of the model is correct?
Explanation: This question assesses understanding of regression in environmental science. The equation y^=12.1−0.18x predicts dissolved oxygen from water temperature. The correct answer (A) properly interprets the slope: for each 1°C increase in temperature, predicted dissolved oxygen decreases by about 0.18 mg/L on average. This uses appropriate language for observational data. Choice B incorrectly claims causation and exact effects at every site. Choice C attempts to extrapolate to 0°C, outside the observed range of 6-24°C. Choice D incorrectly claims independence when negative slope indicates negative association. Choice E completely misinterprets the intercept. While temperature likely does causally affect dissolved oxygen, the regression model itself only describes the observed association within the measured temperature range.
A school counselor collected data from 12 students on the number of hours they studied for a final exam (x, from 1 to 9 hours) and their exam score (y, in points). A least-squares regression line was fit to predict score from study hours: y^=58.2+4.1x. The purpose of this linear model is to summarize the linear association and predict typical exam score from study time within the observed range. Which interpretation of the model is correct?
Explanation: This question tests understanding of slope interpretation in a linear regression model. The regression equation y^=58.2+4.1x models the relationship between study hours and exam scores. The correct interpretation (B) states that for each additional hour studied, the predicted exam score increases by about 4.1 points on average for students similar to those in the data. This properly acknowledges that the slope represents an average association, not a guarantee for individuals. Choice C incorrectly implies causation and exact outcomes for every student. Choice A misinterprets the intercept as a guarantee rather than a prediction. Choice D incorrectly extrapolates to 20 hours, which is far beyond the observed range of 1-9 hours. Choice E confuses the intercept with R-squared. Remember that regression models describe average relationships within the observed data range, not causal effects or guarantees for individuals.
A real estate agent recorded the size of a house (x, in hundreds of square feet) and its selling price (y, in thousands of dollars) for 11 homes in a neighborhood. Sizes ranged from 12 to 28 (i.e., 1200 to 2800 sq ft). A least-squares regression line is y^=95+8.7x. The purpose of this linear model is to summarize the linear relationship and predict typical selling prices for houses within the observed size range. Which interpretation of the model is correct?
Explanation: Interpreting the regression model \hat{y} = 95 + 8.7x for house prices based on size (x in hundreds of sq ft, from 12 to 28) is the key skill. The slope means each 100 sq ft increase is associated with about 8,700 higher predicted price on average within observed sizes. Choice B correctly states this without causal language or extrapolation. Choice C is a distractor, wrongly implying causation from size to price. Choice A dismisses the model due to an unrealistic zero-size intercept, but intercepts can be useful even if extrapolated. Limitations: avoid using the model beyond data, as choice D does for 3500 sq ft. Regression captures linear trends but doesn't account for other variables affecting prices.
A researcher recorded the distance from a city center (x, in miles) and the monthly rent for a one-bedroom apartment (y, in dollars) for 13 apartments, with distances ranging from 1 to 18 miles. A least-squares regression line is y^=1850−42x. The purpose of this linear model is to summarize the linear association and predict typical rents for apartments within the observed distance range. Which interpretation of the model is correct?
Explanation: The skill is interpreting \hat{y} = 1850 - 42x for rent versus distance (x from 1 to 18 miles). The slope shows each mile farther associates with $42 lower predicted rent on average in the range. Choice A is correct, avoiding causation and sticking to data. Choice B distracts by claiming direct causation from distance to rent drop. Choice C misinterprets the intercept as exact for zero miles. Limitation: no extrapolation, unlike choice D to 40 miles. Intercepts estimate averages but may not reflect reality outside data.
A district analyzed 10 schools, recording average class size (x, students per class, from 18 to 34) and average standardized test score (y, points). A least-squares regression line was fit: y^=610−3.5x. The purpose of this linear model is to summarize the linear association and predict typical test score from class size within the observed range. Which interpretation of the model is correct?
Explanation: This question examines proper interpretation of regression in educational policy context. The equation y^=610−3.5x predicts average test scores from average class size. The correct answer (B) properly interprets the slope: for each additional student in average class size, the predicted average test score decreases by about 3.5 points on average. This uses appropriate statistical language avoiding causal claims. Choice A incorrectly implies causation - while smaller classes might cause higher scores, the regression only shows association. Choice C misinterprets the intercept at 0 students per class as meaningful. Choice D wrongly claims class size is the only factor. Choice E incorrectly states that negative slope means zero correlation when it actually indicates negative correlation. Regression models describe associations in observational data but cannot prove causation without proper experimental design.
An environmental scientist models ozone level (y, in ppb) from traffic volume (x, in thousands of cars per day) using data from days with traffic between 10 and 60 (thousand cars). The regression line is y^=18+1.1x. The purpose of the linear model is to describe the association and predict typical ozone levels for traffic volumes in the observed range. Which interpretation of the model is correct?
Explanation: This question tests understanding of slope interpretation in an environmental science context. The regression equation y^=18+1.1x models predicted ozone levels from traffic volume (in thousands of cars), where the slope 1.1 represents the average change in predicted ozone per thousand cars. Choice A correctly states "for each additional 1,000 cars per day, the predicted ozone level increases by about 1.1 ppb, on average." Choice B incorrectly treats the intercept as an actual value rather than a prediction outside the data range. Choice C reverses causation, suggesting ozone causes traffic changes. Choice D claims direct causation from an observational study. Choice E extrapolates to 100 thousand cars, well beyond the observed range of 10-60 thousand. Regression models from observational data describe associations, not causal relationships, and should not be extrapolated beyond their data range.
A researcher studied 13 cars and recorded vehicle weight (x, in thousands of pounds, from 2.4 to 4.8) and highway fuel economy (y, miles per gallon). The least-squares regression line predicting mpg from weight is y^=46.0−5.2x. The purpose of this linear model is to describe the linear association and predict typical fuel economy for weights in the observed range. Which interpretation of the model is correct?
Explanation: This question tests understanding of regression interpretation in an automotive context. The equation y^=46.0−5.2x predicts highway fuel economy from vehicle weight (in thousands of pounds). The correct interpretation (A) states that for each additional 1,000 pounds of weight, predicted highway mpg decreases by about 5.2 on average. This properly uses associative language and acknowledges the average nature of the relationship. Choice B incorrectly implies causation and exact effects. Choice C attempts to interpret the intercept at 0 weight, which is meaningless and far outside the observed range of 2.4-4.8 thousand pounds. Choice D incorrectly claims no relationship when negative slope indicates negative association. Choice E completely misinterprets the intercept. Remember that regression models describe patterns within realistic data ranges, not impossible scenarios like weightless cars.
A manager tracked the number of customers served in an hour (x) and the total tips earned that hour (y, in dollars) for 18 hourly shifts, with x ranging from 12 to 55 customers. A least-squares regression line is y^=8.5+0.62x. The purpose of this linear model is to summarize the linear association and predict typical tips for shifts within the observed range. Which interpretation of the model is correct?
Explanation: This question assesses interpreting \hat{y} = 8.5 + 0.62x for tips from customers served (x from 12 to 55). The slope indicates $0.62 more predicted tips per extra customer on average in the range. Choice A is right, limiting to association and data. Choice B wrongly claims causation for exact increases. Choice C misuses the intercept for zero customers. Key limitation: avoid extrapolation, unlike choice D to 120 customers. Intercepts may not be meaningful alone.
A scientist collected data from 12 batteries on discharge time (x, in hours, from 1.5 to 8.0) and operating temperature (y, in °C, from 28 to 44). The least-squares regression line predicting temperature from discharge time is y^=26.5+2.1x. The purpose of this linear model is to summarize the linear association and predict typical operating temperature for discharge times within the observed range. Which interpretation of the model is correct?
Explanation: This question evaluates understanding of slope interpretation in a technical context. The model y^=26.5+2.1x has a slope of 2.1, meaning for each additional hour of discharge time, the predicted operating temperature increases by 2.1°C on average. Choice A correctly interprets this relationship and appropriately restricts it to the observed range of 1.5-8.0 hours. Choice B incorrectly implies causation and claims an exact temperature rise for every battery. Choice C inappropriately extrapolates to 0 hours discharge time and treats the prediction as accurate outside the data range. Choice D misinterprets the y-intercept as a physical constraint on temperature. Choice E incorrectly claims equal reliability for predictions at 20 hours as within the observed range - extrapolation far beyond observed data is much less reliable. Regression models are tools for understanding patterns within observed data, not for making predictions far outside that range.
A student recorded the number of hours studied (x) and the score on a quiz out of 100 (y) for 12 classmates (hours ranged from 0.5 to 6). A least-squares regression line was fit to predict quiz score from hours studied: y^=52+6.5x. The purpose of this linear model is to summarize the linear association and predict typical quiz scores for study times within the observed range. Which interpretation of the model is correct?
Explanation: This question assesses the skill of interpreting the slope and intercept in a linear regression model for predicting quiz scores from hours studied. The model is y^=52+6.5x, where the slope indicates that for each additional hour studied, the predicted quiz score increases by about 6.5 points on average within the observed range of 0.5 to 6 hours. Choice A correctly captures this associational interpretation without claiming causation or extrapolating beyond the data. A common distractor, like choice C, mistakenly infers causation from the positive slope, assuming that studying causes the score increase, which regression alone cannot prove. Another distractor, choice B, treats the intercept as a literal prediction for x=0, but intercepts often lack real-world meaning outside the data range. A mini-lesson on model limitations: linear regression describes associations but does not imply causation, and predictions should be restricted to the observed range of x to avoid unreliable extrapolations. Always contextualize interpretations to the sample studied, as results may not generalize.
A teacher compared the number of pages a student read in a week (x) with the student's score on a reading quiz (y) for 16 students. Pages ranged from 10 to 80. The least-squares regression line is y^=58+0.35x. The purpose of this linear model is to describe the linear association and predict typical quiz scores for page counts within the observed range. Which interpretation of the model is correct?
Explanation: This question focuses on interpreting y^=58+0.35x, linking pages read (x from 10 to 80) to quiz scores. The slope indicates a 0.35-point increase per extra page on average within the range. Choice A is accurate, emphasizing association and data limits. Distractor C assumes causation, suggesting more reading always boosts scores, but that's not proven. Choice B treats the intercept as exact for zero pages, ignoring variability. Key limitation: no extrapolation, as choice D does to 200 pages. Models like this explain trends but not all variation, and slope isn't r-squared.
An environmental club measured daily high temperature (x, in °F) and the number of bottles of water sold at an outdoor booth (y) for 15 days, with temperatures ranging from 60°F to 92°F. A least-squares regression line was found: y^=−120+4.1x. The purpose of this linear model is to describe the relationship and predict typical sales for temperatures within the observed range. Which interpretation of the model is correct?
Explanation: This question tests the interpretation of a linear regression model relating temperature to water bottle sales, with the equation y^=−120+4.1x. The slope means that for each 1°F increase, predicted sales rise by about 4.1 bottles on average for temperatures between 60°F and 92°F. Choice B is correct as it emphasizes association and limits predictions to observed conditions without causal claims. Choice C is a distractor that incorrectly assumes causation, stating that temperature changes directly cause sales increases, which correlation does not establish. Choice A misinterprets the negative intercept as an exact prediction for 0°F, ignoring that it's outside the data range and may not be meaningful. A key limitation of such models is extrapolation; for instance, predicting at 100°F as in choice E is unreliable because the relationship may not hold beyond observed data. Remember, regression models summarize observed patterns but require caution with intercepts that imply unrealistic scenarios.
A counselor collected data on 10 students: number of absences in a semester (x) and final course percentage (y). Absences ranged from 0 to 12. A least-squares regression line to predict final percentage from absences is y^=93−2.4x. The purpose of this linear model is to summarize the linear association and predict typical final percentages for students with absence counts within the observed range. Which interpretation of the model is correct?
Explanation: The skill here involves correctly interpreting the slope and intercept in a regression model predicting final grades from absences, given by y^=93−2.4x. The negative slope indicates that each additional absence is associated with a 2.4 percentage point decrease in predicted grade on average, within 0 to 12 absences. Choice A accurately reflects this without overstepping into causation or extrapolation. Choice C is a misleading distractor, claiming causation by suggesting reducing absences directly increases grades, but regression shows correlation, not cause. Choice B errs by treating the intercept as a guarantee for zero absences, whereas it's an estimate and actual scores vary. Limitations include avoiding predictions outside the data range, as choice D does by extrapolating to 30 absences, which could be inaccurate if the relationship isn't linear beyond observed values. Overall, these models are tools for description and prediction within limits, not for proving causal effects.
A school nurse recorded the number of minutes a student spent on a treadmill test (x) and the student's heart rate immediately afterward (y, in beats per minute) for 12 students. Times ranged from 3 to 14 minutes. A least-squares regression line is y^=78+5.2x. The purpose of this linear model is to describe the linear association and predict typical heart rates for treadmill times within the observed range. Which interpretation of the model is correct?
Explanation: Interpreting y^=78+5.2x for heart rate after treadmill time (x from 3 to 14 minutes) tests this skill. The slope means each extra minute links to 5.2 bpm higher predicted rate on average in the range. Choice A properly frames it as association. Distractor C assumes causation from time to rate increase. Choice B sees the intercept as proof for zero minutes. Limitation: linearity doesn't guarantee extrapolation, as choice E suggests for 30 minutes. Models predict typical outcomes, not certainties.
A manager tracked advertising spending (x, in hundreds of dollars, from 2 to 25) and weekly sales (y, in thousands of dollars) for 11 weeks. The least-squares regression line predicting sales from ad spending is y^=4.6+0.32x. The purpose of this linear model is to describe the linear association and predict typical weekly sales for ad spending values in the observed range. Which interpretation of the model is correct?
Explanation: This question tests understanding of slope interpretation in a business context. The regression equation y^=4.6+0.32x predicts weekly sales (in thousands) from advertising spending (in hundreds of dollars). The correct interpretation (A) states that spending an additional $100 on advertising is associated with an increase of about 0.32thousand(320) in predicted weekly sales, on average. This properly acknowledges the associative nature and average relationship. Choice B incorrectly claims causation and exact effects. Choice C misinterprets the intercept at x=0 as reliable when the data range is 2-25. Choice D misunderstands what the intercept represents. Choice E incorrectly extrapolates to x=100 (i.e., $10,000), far beyond the observed range. Remember that regression models are descriptive tools for the observed data range, not prescriptive formulas guaranteeing specific outcomes.
A fitness researcher recorded resting heart rate (y, beats per minute) and weekly minutes of aerobic exercise (x, from 0 to 240 minutes) for 14 adults. A scatterplot with the least-squares regression line is shown, with fitted equation y^=78.0−0.06x. The purpose of the linear model is to describe the linear relationship and predict typical heart rate from exercise time within the observed range. Which interpretation of the model is correct?
Explanation: This question examines interpretation of a health-related regression model. The equation y^=78.0−0.06x predicts resting heart rate from weekly exercise minutes. The correct answer (A) properly interprets the slope: an additional minute of exercise per week is associated with about a 0.06 bpm decrease in predicted resting heart rate, on average. This uses appropriate statistical language avoiding causal claims. Choice B misinterprets the intercept as an exact value for all non-exercisers. Choice C demonstrates the danger of extrapolation - 1000 minutes is far beyond the observed range of 0-240 minutes. Choice D incorrectly implies causation and exact effects for everyone. Choice E confuses the intercept with R-squared. Linear models describe average associations within the observed data range, not causal mechanisms or guarantees for individuals.
A fitness coach records resting heart rate (y, beats per minute) and weekly aerobic exercise time (x, minutes) for 14 clients. A least-squares regression line is fit to predict resting heart rate from exercise time: y^=78.4−0.06x. The purpose of this linear model is to describe the association and predict typical resting heart rate for clients with exercise times similar to those observed. Which interpretation of the model is correct?
Explanation: This question tests interpretation of a regression model with a negative slope in a health context. The equation y^=78.4−0.06x indicates that for each additional minute of weekly exercise, the predicted resting heart rate decreases by 0.06 beats per minute on average. Choice A correctly interprets this relationship using appropriate statistical language and acknowledges the model applies to "clients like those in the data." Choice B incorrectly treats the y-intercept as an exact value, Choice C implies causation, Choice D incorrectly links negative slope to nonlinearity (linear models can have negative slopes), and Choice E confuses the y-intercept value with R-squared percentage. Remember that regression models describe average associations, not individual outcomes or causal effects, and their reliability depends on staying within the observed data range.
A school counselor collects data from 12 students on weekly study time (x, hours) and their quiz score (y, points). A least-squares regression line is fit to predict quiz score from study time, with equation y^=58+3.2x. The purpose of this linear model is to summarize the linear association and predict typical quiz scores for students with study times similar to those observed. Which interpretation of the model is correct?
Explanation: This question tests understanding of slope interpretation in linear regression models. The regression equation y^=58+3.2x has a slope of 3.2, which represents the average change in predicted quiz score for each one-unit increase in study time. Choice A correctly interprets this as "for each additional hour studied per week, the predicted quiz score increases by about 3.2 points, on average." The key phrases "predicted," "on average," and "for students with study times like those in the data" properly acknowledge that this is a statistical model describing typical patterns, not exact values or causal relationships. Choices B and D incorrectly treat predictions as exact values, Choice C incorrectly implies causation, and Choice E confuses the slope with R-squared. Linear regression models describe associations and make predictions about typical values within the observed data range, not exact outcomes or causal effects.