PRAXIS CORE MATH (5733) • DATA INTERPRETATION/REPRESENTATION, STATISTICS, AND PROBABILITY

Distinguish Correlation And Causation — Distinguish correlation from causation.

Understanding why two variables moving together does not prove one causes the other is essential for sound statistical reasoning.

Historical Context & Motivation

The distinction between correlation and causation has been one of the most consequential ideas in the history of scientific reasoning. For centuries, philosophers and scientists observed regularities in nature—the sun rises after the rooster crows, disease follows exposure to foul air—and concluded that one event must cause the other. These errors persisted not because observers were careless, but because the formal tools for separating statistical association from genuine causal mechanisms had not yet been developed. The evolution of this distinction shaped modern experimental science, public health policy, and, critically for PRAXIS test-takers, the way we interpret data in educational and social research.

1748
Hume's Problem of Induction
David Hume argued that we can never directly observe causation—only the constant conjunction of events. His philosophical challenge laid the groundwork for distinguishing mere association from true causal links.
1900
Pearson's Correlation Coefficient
Karl Pearson formalized the correlation coefficient (r), giving researchers a precise numerical measure of linear association. Pearson himself cautioned that correlation does not imply causation.
1950
Hill's Criteria for Causation
Austin Bradford Hill proposed nine criteria—including temporality, dose–response, and biological plausibility—to evaluate whether an observed association might be causal. His framework helped establish the link between smoking and lung cancer.
2000s
Causal Inference & DAGs
Judea Pearl and others developed directed acyclic graphs (DAGs) and the do-calculus, providing a mathematical language for modeling causation. These tools formalized confounding and mediation analysis.

The central question this lesson addresses is deceptively simple: When two variables move together in a data set, what can we legitimately conclude? On the PRAXIS Core Math exam, you will encounter scenarios—often drawn from educational research—where you must determine whether an observed statistical relationship warrants a causal claim or merely reflects correlation. Mastering this distinction is not only a test-preparation skill; it is a foundational competency for any future educator who will read research, assess interventions, and teach students to think critically about evidence.

Core Principles & Definitions

Before examining specific examples, it is essential to establish precise definitions for the concepts that underpin this lesson. Many everyday misinterpretations of data stem from conflating association with explanation—a conflation that carries real consequences in fields ranging from medicine to education policy.

1

Correlation

A statistical relationship in which two variables tend to increase or decrease together (positive correlation) or move in opposite directions (negative correlation). Correlation measures the strength and direction of a linear association but says nothing about why the relationship exists.
2

Causation

A relationship in which a change in one variable (the independent variable) directly produces a change in another (the dependent variable). Establishing causation typically requires controlled experimentation or rigorous causal-inference methods.
3

Confounding Variable

A third variable that influences both the independent and dependent variables, creating a spurious association that can be mistaken for causation. Identifying and controlling for confounders is central to rigorous research design.
4

Lurking Variable

A variable not included in the study that affects the results. Unlike a confounding variable, a lurking variable may be entirely unknown to the researcher, making it a hidden source of misleading correlations.
5

Randomized Controlled Experiment

The gold standard for establishing causation. Subjects are randomly assigned to treatment and control groups, which neutralizes confounders—both known and unknown—by distributing them evenly across groups.
KEY TAKEAWAY
Think of correlation like noticing that every time you carry an umbrella, the streets are wet. The umbrella did not cause the wet streets—rain is the confounding variable that caused both. In the same way, two data trends can march in lockstep without one driving the other. The PRAXIS exam tests whether you can resist the temptation to leap from 'these variables are related' to 'this variable caused that outcome.'

Visual Explanation — Correlation vs. Causation

The diagram below illustrates the three fundamental structural relationships that explain why two variables, X and Y, might be correlated. Understanding these structures is the key to distinguishing genuine causation from mere statistical association.

Three structural explanations for an observed correlation between variables X and Y. Panel 1 shows direct causation (X → Y). Panel 2 shows reverse causation (Y → X). Panel 3 shows a confounding variable Z driving both X and Y, producing a spurious correlation.

When you encounter a PRAXIS question stating that two variables are correlated, your first task is to ask which of these three structural relationships might apply. If the study design is observational—that is, researchers merely measured variables without manipulating them—then any of the three structures could explain the data, and a causal claim is not justified. Only when the study involves random assignment to treatment and control groups can confounders be ruled out with reasonable confidence, allowing a causal interpretation.

Mathematical Framework — Correlation Coefficient

Although the PRAXIS Core Math exam does not require you to compute the Pearson correlation coefficient by hand, understanding its mathematical definition clarifies what correlation actually measures—and, just as importantly, what it does not measure. The coefficient quantifies the linear association between two quantitative variables, nothing more.

PEARSON CORRELATION COEFFICIENT
r = Σ[(xᵢ − x̄)(yᵢ − ȳ)] / √[Σ(xᵢ − x̄)² · Σ(yᵢ − ȳ)²]
Where r ranges from −1 to +1. Values near +1 indicate a strong positive linear relationship; values near −1 indicate a strong negative linear relationship; values near 0 indicate no linear relationship. The formula captures how much xᵢ and yᵢ co-vary relative to their individual variability.

Notice what r does not tell us. It does not reveal whether X causes Y, whether Y causes X, or whether some lurking variable Z drives both. It is purely a measure of how closely the data points cluster around a straight line. A value of r = 0.95 between ice cream sales and drowning rates does not mean ice cream causes drowning—summer heat is the confounding variable that drives both.

COEFFICIENT OF DETERMINATION
r² = (proportion of variance in Y explained by the linear relationship with X)
If r = 0.80, then r² = 0.64, meaning 64% of the variability in Y is accounted for by its linear relationship with X. However, 'accounted for' does not mean 'caused by.' The remaining 36% could include confounders, measurement error, or nonlinear effects.
💡 PRAXIS TIP
When a PRAXIS question provides a correlation coefficient or scatterplot, look for language cues. Words like 'associated with,' 'related to,' or 'tends to' describe correlation. Words like 'causes,' 'leads to,' 'results in,' or 'produces' describe causation. If the study is observational, causation language is not justified regardless of how strong the correlation is.

Study Design — When Can We Claim Causation?

The type of study design determines the strongest conclusion you can draw. On the PRAXIS exam, you will need to evaluate a described study and decide whether a correlation-only statement or a causal statement is warranted. The diagram below maps out the key decision points.

The flowchart shows the critical decision point: random assignment. When participants are randomly assigned to groups, confounders are distributed evenly, allowing causal claims. Without random assignment, only correlation can be stated.
Comparison of observational and experimental study designs
FeatureObservational StudyRandomized Experiment
Random AssignmentNo — groups are self-selected or pre-existingYes — researcher assigns subjects to groups
ConfoundersMay be present and unmeasuredDistributed evenly across groups by randomization
Strongest ConclusionCorrelation (association)Causation (if well-designed and replicated)
ExampleSurvey finds students who eat breakfast score higherStudents randomly assigned to receive breakfast; their scores are compared to control

Worked Example — Evaluating a Research Claim

Consider the following PRAXIS-style scenario: A school district reports that students who participated in an after-school tutoring program had, on average, 15% higher scores on the state math assessment compared to students who did not participate. The district claims that the tutoring program caused the improvement. Evaluate this claim.

Evaluating the Tutoring Program Claim
1
Step 1 — Identify the VariablesThe independent variable is participation in the tutoring program (yes or no). The dependent variable is the score on the state math assessment. The reported finding is that these two variables are positively associated.
X = tutoring participation; Y = math score; r is positive.
2
Step 2 — Determine the Study DesignThe scenario says students 'participated' in the program—they were not randomly assigned. Students (or their parents) chose to participate, making this an observational study, not a randomized experiment.
Study type: observational (self-selected groups).
3
Step 3 — Identify Potential ConfoundersStudents who voluntarily attend after-school tutoring may differ systematically from those who do not. They may be more motivated, come from families that prioritize education, or have greater access to transportation. Any of these confounding variables could independently predict higher test scores.
Confounders: motivation, family support, socioeconomic status.
4
Step 4 — Evaluate the Causal ClaimBecause the study is observational, the 15% score difference could be driven by the confounders identified in Step 3 rather than by the tutoring program itself. Without random assignment, we cannot rule out these alternative explanations.
The causal claim is NOT justified.
5
Step 5 — State the Correct ConclusionThe data show a correlation between tutoring participation and higher test scores. A valid conclusion would be: 'Participation in the tutoring program is associated with higher math scores.' To establish causation, the district would need to randomly assign students to tutoring and control groups and compare outcomes.
Correct answer: The data show correlation, not causation.

Common Errors & Reasoning Pitfalls

On the PRAXIS exam, incorrect answer choices are often designed to exploit common reasoning errors. Recognizing these pitfalls is just as important as understanding the correct framework. The table below catalogues the most frequent mistakes and how to avoid them.

Common reasoning errors on correlation vs. causation questions
Reasoning ErrorDescriptionHow to Avoid It
Post hoc fallacyAssuming that because event B followed event A, A caused B. ('After this, therefore because of this.')Temporal sequence is necessary but not sufficient for causation. Demand evidence of mechanism and controlled comparison.
Ignoring confoundersTreating a strong correlation as proof of causation without considering third variables that could produce the association.Always ask: 'Could a lurking variable explain this relationship?' before accepting a causal claim.
Reverse causationAssuming X causes Y when, in fact, Y causes X. E.g., concluding that hospital stays cause illness.Examine whether the presumed direction of causation is logically coherent and supported by temporal evidence.
Over-generalizing from rInterpreting a high correlation coefficient as evidence of causation. r = 0.99 between two variables does not establish causation.Remember: r measures the strength of linear association, not the presence of a causal mechanism.
Denying all relationshipsOvercorrecting by claiming that because a study is observational, the variables have no real relationship at all.Correlation is real—it just doesn't prove causation. An observed association is still informative and worth reporting.
KEY TAKEAWAY
Think of correlation and causation like a detective's evidence and verdict. A fingerprint at the crime scene (correlation) is evidence, but it does not prove guilt (causation). The fingerprint could have been left before the crime (temporal confound), planted by someone else (lurking variable), or left by the victim who turned out to be involved differently than assumed (reverse causation). A detective needs a rigorous investigation—an experiment—to establish what actually happened.

Connection to Advanced Statistical Reasoning

While the PRAXIS Core Math exam assesses fundamental understanding of correlation versus causation, the distinction connects to more advanced topics in statistics and research methodology that you may encounter as a practicing educator. Understanding where the PRAXIS-level knowledge fits within the broader statistical landscape will deepen your conceptual fluency and prepare you for graduate-level coursework in education research.

PRAXIS concepts and their advanced statistical extensions
PRAXIS-Level ConceptAdvanced Extension
Correlation coefficient (r) measures linear associationMultiple regression controls for several variables simultaneously to isolate partial correlations and estimate adjusted effects
Observational studies cannot prove causationQuasi-experimental designs (difference-in-differences, regression discontinuity) can strengthen causal claims in settings where randomization is impractical
Confounding variables create spurious correlationsDirected acyclic graphs (DAGs) formalize confounding, mediation, and collider bias, guiding variable selection in research models
Randomized experiments are the gold standardMeta-analyses aggregate results across multiple experiments to estimate effect sizes with greater precision and generalizability

As future educators, you will regularly encounter research claims about instructional practices, curricular interventions, and assessment strategies. Being able to evaluate whether a study's design supports a causal claim—or merely an associational one—will empower you to make evidence-based decisions in your classroom and to teach your own students the critical thinking skills that underpin scientific literacy.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher finds that students who listen to classical music tend to have higher GPAs. The researcher concludes that listening to classical music improves academic performance. What is the primary flaw in this reasoning?
PROBLEM 2BASIC CALCULATION
A study reports a Pearson correlation coefficient of r = 0.85 between hours spent on homework and exam scores. Calculate r² and interpret it. Does this value prove that doing more homework causes higher scores?
PROBLEM 3INTERMEDIATE
A school board reviews two studies about a reading intervention. Study A: Teachers voluntarily adopted the intervention; their students showed 20% gains over non-adopting teachers' students. Study B: 50 classrooms were randomly assigned to use the intervention or standard instruction; intervention classrooms showed a 12% gain. Which study provides stronger evidence that the intervention caused improved reading scores, and why?
PROBLEM 4APPLIED
A newspaper headline reads: 'States with more teachers per student have lower crime rates.' A reader concludes that hiring more teachers would reduce crime. Identify at least two confounding variables that could explain this correlation, and describe what kind of study would be needed to support the causal claim.
PROBLEM 5CRITICAL THINKING
A colleague argues: 'We can never truly prove causation because there could always be an unmeasured confounding variable, even in randomized experiments. Therefore, the distinction between correlation and causation is meaningless.' Critically evaluate this argument. In what sense is the colleague correct, and in what sense does the argument go too far?

Lesson Summary

Correlation describes a statistical relationship in which two variables move together, measured by the Pearson correlation coefficient (r), which ranges from −1 to +1. Causation means that a change in one variable directly produces a change in another. The critical insight is that correlation does not imply causation because confounding variables, reverse causation, and lurking variables can all produce associations between variables that have no direct causal link.

To establish causation, a study must use random assignment to treatment and control groups, which distributes confounders evenly and isolates the effect of the independent variable. Observational studies—no matter how large or how strong the correlation—can only demonstrate association, not causation. On the PRAXIS Core Math exam, look for keywords: 'associated with' signals correlation; 'causes' or 'leads to' signals a causal claim that must be supported by experimental evidence. As future educators, mastering this distinction equips you to evaluate research, make evidence-based instructional decisions, and model critical thinking for your students.

Varsity Tutors • PRAXIS Core Math (5733) • Distinguish Correlation And Causation