AP STATISTICS • EXPLORING TWO-VARIABLE DATA

Linear Regression Models

Quantifying the linear relationship between two variables to make predictions grounded in data.

Historical Context & Motivation

The desire to fit a straight line to data is far older than the discipline of statistics itself. In the late eighteenth century, astronomers faced the practical challenge of combining multiple, slightly discrepant measurements of planetary orbits into a single best estimate. The method of least squares emerged from this need, providing a principled criterion for choosing one line over infinitely many alternatives. Over the next two centuries, the technique was refined into the modern linear regression model — arguably the most widely used statistical tool in science, economics, medicine, and public policy.

1805
Legendre Publishes Least Squares
Adrien-Marie Legendre published Nouvelles Méthodes pour la Détermination des Orbites des Comètes, introducing the least-squares criterion as a practical fitting technique for astronomical data.
1809
Gauss Provides Probabilistic Foundation
Carl Friedrich Gauss demonstrated that least squares yields the most probable parameter estimates when errors follow a normal distribution, linking regression to probability theory.
1885
Galton Coins 'Regression'
Francis Galton studied the heights of parents and children, noting that extreme parental heights tended to produce offspring closer to the population mean — a phenomenon he called regression toward the mean.
1897
Pearson Formalizes Correlation
Karl Pearson developed the product-moment correlation coefficient r, giving researchers a dimensionless measure of linear association that connects directly to the slope of the regression line.
1922
Fisher Establishes Modern Inference
Ronald Fisher formalized the sampling distributions of regression coefficients, enabling hypothesis tests and confidence intervals that remain central to the AP Statistics curriculum.

The central question that drove all of this work remains the same today: given a cloud of data points in two dimensions, how do we find the single straight line that best summarizes the relationship between the explanatory variable x and the response variable y? And once we have that line, how confidently can we use it to predict new outcomes or infer the nature of the underlying process?

Core Principles & Definitions

A linear regression model rests on a small number of foundational ideas. Before diving into formulas, you should internalize the conceptual framework: what the model assumes, what quantities it estimates, and what yardstick it uses to judge a good fit. The four concepts below form the backbone of every regression analysis you will encounter on the AP Statistics exam.

1

Least-Squares Criterion

The least-squares regression line (LSRL) is the unique line that minimizes the sum of the squared vertical distances (residuals) from each data point to the line. Squaring ensures that positive and negative deviations do not cancel.
2

Slope & Intercept

The regression equation ŷ = a + bx has two parameters: the y-intercept a (predicted y when x = 0) and the slope b (predicted change in y for each one-unit increase in x).
3

Residuals

A residual is the difference between an observed y-value and the predicted value ŷ: residual = y − ŷ. Residual analysis is the primary diagnostic for checking model adequacy.
4

Coefficient of Determination (r²)

The value tells you what fraction of the total variability in y is accounted for by the linear relationship with x. It ranges from 0 (no linear fit) to 1 (perfect linear fit).
KEY TAKEAWAY
Think of the LSRL as a GPS route that minimizes total travel time. Among all possible straight-line routes through the data, the least-squares line is the one that minimizes the total 'cost' — measured as the sum of squared residuals. Just as a GPS recalculates when conditions change, the regression line updates its slope and intercept whenever new data arrive, always targeting the lowest possible total squared error.

Visual Explanation — The Scatterplot & LSRL

A scatterplot with the regression line overlaid is the single most important visual in any regression analysis. It allows you to assess the direction and strength of the relationship at a glance, spot potential outliers, and detect non-linear patterns that might violate the model's assumptions. The diagram below illustrates a dataset with a moderately strong positive linear association together with the LSRL and several key features referenced throughout this lesson.

Each violet dot is an observed data point. The cyan line is the LSRL. The pink dashed segment shows one residual (y − ŷ). The amber point marks (x̄, ȳ), through which every LSRL passes.

Notice that the regression line passes through the point (x̄, ȳ). This is not a coincidence — it is a mathematical consequence of the least-squares derivation. When you interpret a scatterplot on the AP exam, start by noting three things: the direction (positive or negative slope), the form (linear, curved, or clustered), and the strength (how tightly the points hug the line). These three descriptors — direction, form, and strength — are the language of scatterplot interpretation expected in free-response questions.

Mathematical Framework

The formulas below define the LSRL and its key summary statistics. You are expected to interpret these quantities on the AP exam; while calculator output typically provides the numerical values, understanding where they come from gives you the reasoning power needed for free-response explanations.

REGRESSION EQUATION
ŷ = a + bx
ŷ is the predicted value of y. The parameter b is the slope and a is the y-intercept. Some textbooks write ŷ = b₀ + b₁x.
SLOPE
b = r × (s_y / s_x)
Here r is the correlation coefficient, sy is the standard deviation of y, and sx is the standard deviation of x. The slope inherits the sign of r.
Y-INTERCEPT
a = ȳ − b × x̄
This formula guarantees that the point (x̄, ȳ) lies on the regression line. In context, the intercept is the predicted y when x = 0, though this prediction is often outside the range of observed data and may lack practical meaning.
COEFFICIENT OF DETERMINATION
r² = 1 − (SSE / SST)
SSE = Σ(y − ŷ)² is the sum of squared residuals (error). SST = Σ(y − ȳ)² is the total sum of squares. The ratio SSE/SST measures the fraction of variability not explained by the model, so r² gives the fraction that is explained.
🔢 Calculator Tip
On a TI-83/84, run LinReg(a+bx) from the STAT → CALC menu. The output gives a, b, r², and r. Make sure Diagnostics are turned ON (in the CATALOG) to display r and r². The AP exam expects you to read, interpret, and report these values in context.

Residual Analysis & Model Diagnostics

Fitting a line is only half the story. The other half is checking whether a linear model is appropriate. The primary diagnostic tool is the residual plot — a scatterplot of residuals (y − ŷ) on the vertical axis against the explanatory variable x (or the predicted values ŷ) on the horizontal axis. If the linear model is a good fit, the residual plot should show a random scatter with no obvious pattern. Systematic curvature, fanning, or clusters indicate that the linear model is inadequate and a different model form or transformation should be considered.

Top-left: a healthy residual plot showing random scatter. Top-right: a curved pattern indicates a non-linear relationship. Bottom-left: fanning (heteroscedasticity) means the spread of residuals increases with x. Bottom-right: key properties of residuals.

On the AP exam, you will often be asked to determine whether a linear model is appropriate based on a residual plot. The decision rule is straightforward: if the residual plot exhibits no discernible pattern and the residuals are roughly symmetric about the horizontal zero line, a linear model is reasonable. Any systematic curvature, clustering, or change in spread suggests the model needs revision. Additionally, pay attention to influential points — observations with unusually large leverage (extreme x-values) or large residuals (outliers in the y-direction). A point is influential if removing it substantially changes the slope or intercept of the regression line.

📝 AP Exam Vocabulary
An outlier is a point whose residual is much larger in absolute value than the other residuals. A high-leverage point has an x-value far from x̄. A point that is both high-leverage and an outlier is almost always influential — it pulls the regression line toward itself.

Worked Example

A researcher collects data on the number of hours studied (x) and exam scores (y) for 8 students. The summary statistics are: x̄ = 5.0, ȳ = 74.0, sx = 2.0, sy = 8.0, and r = 0.85. Find the equation of the LSRL, interpret the slope and r² in context, and predict the exam score for a student who studies 7 hours.

Finding and Interpreting the LSRL
1
Step 1 — Calculate the Slope bUsing the formula b = r × (sy / sx), substitute: b = 0.85 × (8.0 / 2.0) = 0.85 × 4.0
b = 3.40
2
Step 2 — Calculate the Y-Intercept aUsing a = ȳ − b × x̄, substitute: a = 74.0 − 3.40 × 5.0 = 74.0 − 17.0
a = 57.0
3
Step 3 — Write the Regression EquationSubstituting a and b into the LSRL form:
ŷ = 57.0 + 3.40x
4
Step 4 — Interpret the Slope in ContextFor each additional hour spent studying, the model predicts that a student's exam score increases by 3.40 points, on average. The AP exam requires you to mention the units of both variables and the phrase 'predicted' or 'on average' to signal that this is a model-based statement, not a certainty.
5
Step 5 — Compute and Interpret r²r² = (0.85)² = 0.7225. In context: approximately 72.25% of the variability in exam scores is explained by the linear relationship with hours studied.
r² ≈ 0.7225 (72.25%)
6
Step 6 — Make a PredictionFor a student who studies x = 7 hours: ŷ = 57.0 + 3.40(7) = 57.0 + 23.8. Since x = 7 is within the observed range of x-values, this is an interpolation and is considered a reliable prediction.
ŷ = 80.8 points
⚠️ Extrapolation Warning
If the student studied x = 15 hours (well beyond the observed data range of roughly 1–9 hours), the model would predict ŷ = 57.0 + 3.40(15) = 108.0 — an impossible exam score if the maximum is 100. Extrapolation — using the model outside the range of observed x-values — is unreliable because there is no evidence the linear trend continues beyond the data.

Interpretation Guidelines & Common Pitfalls

On the AP exam, clear and precise interpretation of regression output earns or loses you credit on free-response questions. The table below summarizes correct versus incorrect interpretive language for the most frequently tested quantities. Memorizing these templates will help you write concise, rubric-aligned responses under time pressure.

Interpretation templates for key regression quantities
QuantityCorrect Interpretation TemplateCommon Error
Slope (b)For each additional [unit of x], the predicted [y] increases/decreases by [b] [units of y], on average.Omitting 'predicted' or 'on average,' implying causation, or forgetting context (variable names and units).
Intercept (a)When [x] = 0, the predicted [y] is [a]. (Add: 'This may not have practical meaning if x = 0 is outside the data range.')Interpreting the intercept as meaningful when x = 0 is nonsensical (e.g., 0 hours of sunlight).
Approximately [r² × 100]% of the variability in [y] is explained by the linear relationship with [x].Saying 'r² percent of [y] is caused by [x]' or confusing r² with r.
r (correlation)There is a [strong/moderate/weak], [positive/negative], linear association between [x] and [y].Using 'relationship' instead of 'linear association,' implying causation, or describing a non-linear pattern as strong because |r| is large.
ResidualThe actual [y] was [residual value] [units] above/below the value predicted by the model.Confusing the sign: positive residual means the observed y is above the predicted value, not below.
CORRELATION ≠ CAUSATION
A strong linear regression between ice cream sales and drowning deaths does not mean ice cream causes drowning. Both are driven by a lurking variable — temperature. Regression describes association; establishing causation requires a well-designed randomized experiment. On the AP exam, avoid causal language unless the data come from an experiment with random assignment.

Connection to Inference for Regression

Everything covered so far falls under descriptive regression — summarizing a dataset. The AP Statistics curriculum also includes inference for regression, which asks whether the observed linear association in a sample provides convincing evidence of a true linear relationship in the population. This inferential step involves hypothesis tests and confidence intervals for the population slope β, and it requires additional conditions beyond what we have discussed.

Descriptive regression vs. inference for regression
FeatureDescriptive Regression (This Lesson)Inference for Regression (Later Unit)
GoalSummarize the relationship in the observed dataDetermine whether the relationship exists in the population
Key Statisticb (sample slope), r, r²t = b / SE(b), with df = n − 2
ConditionsLinearity (check residual plot)Linearity, independence, normality of residuals, equal variance (LINE)
OutputEquation: ŷ = a + bx; r²p-value, confidence interval for β
HypothesisNot applicableH₀: β = 0 (no linear relationship) vs. Hₐ: β ≠ 0

The bridge between the two topics is the acronym LINE: Linearity, Independence, Normality of residuals, and Equal variance (homoscedasticity). In descriptive regression, you focus primarily on the L; in inferential regression, all four conditions must be verified. Mastering residual analysis now will pay dividends when you encounter t-tests and confidence intervals for the slope later in the course.

Practice Problems

1
A least-squares regression line is fit to a scatterplot of y versus x. Which of the following is always true about the residuals from this regression?
2
For a dataset, x̄ = 10, ȳ = 50, sx = 4, sy = 12, and r = −0.60. What is the equation of the least-squares regression line?
3
A regression of test score (y) on hours of sleep the night before (x) yields ŷ = 40 + 5x with r² = 0.64. A student who slept 8 hours scored 76 on the test. What is this student's residual, and what does it indicate?
PROBLEM 4APPLIED
An environmental scientist collects data on average daily temperature (°F) and daily electricity consumption (kWh) for 30 summer days. Computer output gives the following: Predictor Coef SE Coef T P Constant −50.2 18.1 −2.77 0.010 Temperature 6.80 0.21 32.38 <0.001 S = 14.3 R-Sq = 97.4% R-Sq(adj) = 97.3% (a) Write the equation of the least-squares regression line. Define any variables used. (b) Interpret the slope in the context of this problem. (c) Interpret r² in context. (d) The scientist wants to predict electricity consumption on a day with a high of 120°F, which is outside the observed temperature range of 72°F to 105°F. Should the scientist use this regression model for that prediction? Explain.
PROBLEM 5CRITICAL THINKING
A student fits a least-squares regression line to a dataset and obtains r² = 0.92. The student concludes: 'Since r² is very high, the linear model is an excellent fit for these data, and x causes changes in y.' (a) Identify two specific statistical errors or unwarranted claims in the student's conclusion. (b) Describe one graphical method the student could use to assess whether the linear model is truly appropriate, and explain what the student should look for. (c) Give a real-world example of two variables that could have r² ≈ 0.92 but where x clearly does not cause y. Explain the role of a lurking variable.

Summary

A linear regression model summarizes the relationship between an explanatory variable x and a response variable y using the equation ŷ = a + bx, where the slope b equals r × (s_y / s_x) and the intercept a equals ȳ − b × x̄. The line always passes through the point (x̄, ȳ) and is chosen to minimize the sum of squared residuals, where each residual is y − ŷ.

The coefficient of determination r² measures the fraction of variability in y explained by the linear model. Always check a residual plot for random scatter before trusting the model — a high r² alone is not sufficient. Avoid extrapolation (predicting beyond the observed data range) and never infer causation from observational regression. Interpret every statistic — slope, intercept, r², and residuals — in the context of the data using the variable names and units.

Varsity Tutors • AP Statistics • Linear Regression Models