AP STATISTICS • EXPLORING TWO-VARIABLE DATA

Representing Two Categorical Variables

Two-way tables and segmented bar charts reveal the hidden relationships between categorical variables.

Historical Context & Motivation

Long before the formal machinery of statistical inference was developed, researchers recognized that understanding relationships between categorical attributes—such as occupation and literacy, or treatment type and patient outcome—required more than simple one-variable frequency counts. The challenge was clear: when two categorical variables each have multiple levels, summarizing every possible combination of categories demands a structured, tabular approach. The two-way table (also called a contingency table) emerged as the foundational tool for organizing such data, and its development parallels the rise of quantitative social science.

1693
Halley's Mortality Table
Edmond Halley published a pioneering life table for the city of Breslau, cross-classifying age groups with survival counts—one of the earliest uses of tabular data to study a categorical relationship.
1900
Pearson's Chi-Square Test
Karl Pearson introduced the chi-square goodness-of-fit test, providing the first rigorous method for assessing whether observed frequencies in a two-way table differ significantly from expected frequencies under independence.
1922
Fisher and the Contingency Table
Ronald A. Fisher refined the analysis of contingency tables with exact tests and maximum-likelihood methods, establishing the theoretical framework that AP Statistics students still use today.
1977
Exploratory Data Analysis (EDA)
John Tukey's emphasis on visual displays of data led to the widespread adoption of segmented bar charts, mosaic plots, and other graphical representations of two-way tables.

The central question this topic addresses is deceptively simple: Is there a relationship between two categorical variables, and how can we display and quantify that relationship? To answer it, we need tools that organize raw counts into meaningful summaries and graphical displays that make patterns—and their absence—visually apparent. This section of the AP Statistics curriculum equips you with exactly those tools.

Core Principles & Definitions

Before constructing any table or chart, it is essential to have a firm grasp of the key vocabulary. A categorical variable places each individual into one of several groups or categories; examples include political affiliation, blood type, or preferred mode of transportation. When we observe two categorical variables for each individual in a dataset, we can organize the counts into a two-way table (also called a contingency table), where rows represent the categories of one variable and columns represent the categories of the other.

1

Joint Frequency

The count (or proportion) of individuals who fall into a specific combination of categories from both variables. Each interior cell of a two-way table is a joint frequency.
2

Marginal Frequency

The total count (or proportion) for a single category of one variable, summed across all categories of the other variable. These appear in the 'margins' (last row and last column) of the table.
3

Conditional Distribution

The distribution of one variable among only those individuals who belong to a specific category of the other variable. Comparing conditional distributions is the key to detecting an association.
4

Marginal Distribution

The overall distribution of a single variable, ignoring the other variable entirely. Derived from the marginal frequencies and useful as a baseline for comparison.
5

Association vs. Independence

Two categorical variables are associated if the conditional distribution of one variable changes across the categories of the other. If the conditional distributions are (approximately) identical, the variables are independent.
KEY TAKEAWAY
Think of a two-way table like a spreadsheet pivot table: it takes a long list of individual records, each tagged with two labels, and collapses them into a compact grid of counts. The marginal totals are the row and column sums that appear along the edges, much like the grand totals in an accounting ledger. Comparing the conditional distributions across rows (or columns) is analogous to A/B testing in tech—if the outcome proportions differ meaningfully between groups, an association may exist.

Visual Explanation: The Two-Way Table

The diagram below presents a two-way table for a hypothetical survey of 400 college students cross-classified by class year (Freshman, Sophomore, Junior, Senior) and preferred study method (Alone, Group, Online). Joint frequencies appear in the interior cells; marginal frequencies are displayed along the right edge and bottom row. Observe how the marginal totals for each row sum to the row total, and the marginal totals for each column sum to the column total; both grand totals equal 400.

A two-way table cross-classifying 400 students by class year and preferred study method. Interior cells show joint frequencies; the dashed-line borders mark marginal frequencies along the bottom row and right column.

Notice the structural anatomy of the table. The intersection of the Freshman row and the Alone column yields the joint frequency 50, meaning 50 freshmen in the sample prefer studying alone. The row total for Freshmen (120) is the marginal frequency of the Freshman category, obtained by summing 50 + 30 + 40. Similarly, the column total for Alone (140) sums 50 + 35 + 30 + 25. The grand total of 400 in the bottom-right corner serves as a quick check: every row total and every column total must each sum to 400.

Mathematical Framework: Proportions & Distributions

Raw counts are informative, but proportions allow us to make meaningful comparisons—especially when row or column totals are unequal. There are three types of proportions commonly extracted from a two-way table: joint relative frequencies, marginal relative frequencies, and conditional relative frequencies. Each answers a fundamentally different question.

JOINT RELATIVE FREQUENCY
Joint proportion = (cell count) / (grand total)
For example, the joint proportion of Freshmen who study Alone = 50 / 400 = 0.125 (or 12.5%). This answers: "What fraction of all students are both Freshmen and prefer Alone?"
MARGINAL RELATIVE FREQUENCY
Marginal proportion = (row or column total) / (grand total)
For example, the marginal proportion of Freshmen = 120 / 400 = 0.30. This answers: "What fraction of all students are Freshmen, regardless of study method?"
CONDITIONAL RELATIVE FREQUENCY (ROW-BASED)
Conditional proportion = (cell count) / (row total)
For example, among Freshmen only, the conditional proportion who study Alone = 50 / 120 ≈ 0.417 (41.7%). This answers: "Of the Freshmen, what fraction prefer studying Alone?" Comparing these across rows is the primary method for detecting association.
CONDITIONAL RELATIVE FREQUENCY (COLUMN-BASED)
Conditional proportion = (cell count) / (column total)
For example, among those who study Alone, the conditional proportion who are Freshmen = 50 / 140 ≈ 0.357. This answers: "Of those who study Alone, what fraction are Freshmen?"
📝 AP Exam Tip
Free-response questions frequently ask you to compute conditional proportions and then compare them to determine whether an association exists. Always specify clearly what the denominator represents (e.g., "among Freshmen" or "among those who study Alone") and state your comparison in context. Failure to reference the conditioning group is a common point deduction.

Graphical Displays: Segmented & Side-by-Side Bar Charts

While the two-way table is the computational backbone, graphical representations make associations (or their absence) immediately visible. The AP Statistics curriculum emphasizes two primary chart types for two categorical variables: segmented (stacked) bar charts and side-by-side bar charts. A segmented bar chart displays the conditional distribution of one variable within each category of the other by stacking colored segments to fill each bar to 100%. If the segment proportions are roughly the same across all bars, the variables appear independent; if the proportions shift noticeably, an association is suggested.

Segmented bar chart showing the conditional distribution of study method within each class year. The varying segment sizes across bars suggest an association: the proportion preferring Group study increases from Freshman (25.0%) to Senior (50.0%), while the Alone proportion decreases.

The chart makes the association visually apparent. The cyan segment (Alone) shrinks from 41.7% for Freshmen to 27.8% for Seniors, while the violet segment (Group) grows from 25.0% to 50.0%. If the two variables were independent—meaning class year had no relationship with study preference—all four bars would have essentially identical segment proportions. The visible differences suggest that as students advance through college, they increasingly prefer group study. A side-by-side bar chart would display the same information by placing separate bars for each study method next to one another within each class-year group, making it easier to compare absolute heights but harder to see proportional composition.

Worked Example

A health researcher surveys 500 adults about their exercise frequency (Regular, Occasional, None) and whether they report high stress levels (Yes, No). The data are shown in the table below. Determine whether there appears to be an association between exercise frequency and high stress.

Two-way table: Exercise Frequency × High Stress
High Stress: YesHigh Stress: NoTotal
Regular40160200
Occasional6090150
None9060150
Total190310500
Detecting Association via Conditional Distributions
1
Step 1 — Identify the VariablesThe explanatory variable is exercise frequency (Regular, Occasional, None) and the response variable is high stress (Yes, No). We will compute the conditional distribution of high stress for each level of exercise frequency.
2
Step 2 — Compute Conditional Proportions of High StressFor each exercise group, divide the number reporting high stress by the row total. Regular: 40 / 200 = 0.20 (20%). Occasional: 60 / 150 = 0.40 (40%). None: 90 / 150 = 0.60 (60%).
Regular → 20%, Occasional → 40%, None → 60%
3
Step 3 — Compare the Conditional DistributionsThe conditional proportion of high stress differs substantially across the three exercise groups: it triples from 20% (Regular) to 60% (None). If exercise frequency and high stress were independent, all three conditional proportions would be approximately equal to the marginal proportion of high stress, which is 190 / 500 = 0.38 (38%).
4
Step 4 — State Your Conclusion in ContextBecause the conditional proportions of high stress vary considerably across exercise-frequency groups (20%, 40%, 60%), there appears to be an association between exercise frequency and high stress. Adults who exercise regularly report high stress at a much lower rate than those who exercise occasionally or not at all.
Conclusion: There is an association between exercise frequency and reporting high stress.
⚠️ Caution: Association ≠ Causation
Observing an association in a two-way table does not establish a causal relationship. The data above come from a survey (observational study), not a randomized experiment, so confounding variables—such as income, occupation, or prior health conditions—may partly or fully explain the observed pattern.

Strengths & Limitations of Graphical Displays

Different graphical displays of two categorical variables emphasize different aspects of the data. Choosing the right display depends on whether your goal is to compare conditional proportions, compare counts, or show the overall composition. The table below summarizes the strengths and limitations of the most common options.

Comparison of common displays for two categorical variables
Display TypeBest ForLimitations
Two-Way TablePrecise counts and proportions; computing conditional distributions; the essential starting point for any analysisNot a visual display; patterns can be hard to spot in large tables with many categories
Segmented Bar ChartComparing conditional distributions across groups; detecting association at a glanceDifficult to compare middle segments accurately; does not show sample sizes unless annotated
Side-by-Side Bar ChartComparing absolute counts or proportions for specific categories; easy to read individual bar heightsHarder to see the full conditional distribution within each group; chart can become cluttered with many categories
Mosaic PlotShowing both conditional proportions and relative group sizes in a single displayLess intuitive to read; not explicitly tested on the AP exam, but appears occasionally
KEY TAKEAWAY
Think of choosing a chart type like choosing the right lens for a camera. A segmented bar chart is a wide-angle lens: it captures the full proportional composition in one frame. A side-by-side bar chart is more like a zoom lens: it isolates individual categories for precise comparison. On the AP exam, the segmented bar chart is the most frequently tested display because it directly visualizes the conditional distributions you need to discuss association.

Connection to Inference: The Chi-Square Test

Everything covered so far—two-way tables, conditional distributions, and segmented bar charts—belongs to the exploratory data analysis (EDA) stage of statistics. We can describe patterns and note apparent associations, but we cannot yet determine whether those patterns are statistically significant or could have arisen by chance. That is the role of inferential statistics, specifically the chi-square test for independence, which you will encounter later in the AP Statistics curriculum.

Descriptive analysis (this lesson) vs. inferential chi-square test (later unit)
FeatureDescriptive (This Lesson)Inferential (Chi-Square Test)
GoalSummarize and display the relationship between two categorical variables in a sampleDetermine whether a relationship observed in a sample provides convincing evidence of a relationship in the population
ToolsTwo-way tables, segmented bar charts, conditional proportionsExpected counts, chi-square statistic (χ²), p-values, degrees of freedom
Output"There appears to be an association...""There is statistically significant evidence of an association (p < 0.05)..."
ScopeLimited to the sample at handGeneralizes from sample to population under stated conditions

Mastering the descriptive skills in this lesson is essential groundwork for the chi-square test. The expected counts used in that test are derived from the marginal totals you already know how to compute, and the test statistic measures how far the observed joint frequencies deviate from what we would expect under independence. In other words, the chi-square test formalizes exactly the kind of comparison you perform when you eyeball a segmented bar chart and decide whether the segment proportions look "different enough" to matter.

Practice Problems

1
A researcher constructs a two-way table of political party affiliation (Democrat, Republican, Independent) and opinion on a policy proposal (Support, Oppose). To determine whether there is an association between party affiliation and opinion, the researcher should compare which of the following?
2
In a survey of 600 employees, 180 work remotely and 420 work on-site. Among remote workers, 108 report high job satisfaction. Among on-site workers, 210 report high job satisfaction. What is the conditional proportion of high job satisfaction among remote workers?
3
A segmented bar chart shows the conditional distribution of transportation mode (Car, Bus, Bike) for two neighborhoods, A and B. In Neighborhood A, the segments are approximately Car 60%, Bus 25%, Bike 15%. In Neighborhood B, the segments are approximately Car 35%, Bus 40%, Bike 25%. Which of the following conclusions is best supported by this display?
PROBLEM 4APPLIED
A university admissions office collects data on 1,200 applicants classified by admission decision (Admitted, Denied) and intended major division (STEM, Humanities, Social Sciences). The data are shown below. | | Admitted | Denied | Total | |---|---|---|---| | STEM | 200 | 300 | 500 | | Humanities | 150 | 150 | 300 | | Social Sciences | 180 | 220 | 400 | | Total | 530 | 670 | 1200 | (a) Calculate the conditional proportion of admission for each major division. (b) Describe what these proportions reveal about the relationship between major division and admission decision. (c) A segmented bar chart is constructed with bars for each major division and segments representing Admitted and Denied. Describe what you would expect to see in the chart if the two variables were independent. (d) A student says, "STEM applicants are discriminated against because they have the lowest admission rate." Identify a flaw in this reasoning.
PROBLEM 5CRITICAL THINKING
A hospital compares two treatments (A and B) for kidney stones. The data, broken down by stone size, are: Small stones: Treatment A succeeds in 81 of 87 cases (93%). Treatment B succeeds in 234 of 270 cases (87%). Large stones: Treatment A succeeds in 192 of 263 cases (73%). Treatment B succeeds in 55 of 80 cases (69%). Overall: Treatment A succeeds in 273 of 350 cases (78%). Treatment B succeeds in 289 of 350 cases (83%). (a) Verify that Treatment A has a higher success rate than Treatment B for both small stones and large stones separately. (b) Explain how it is possible that Treatment B has a higher overall success rate despite Treatment A being better within each subgroup. (c) What is this phenomenon called, and what is the lurking variable responsible? (d) A hospital administrator wants to recommend one treatment. Using the data, explain which treatment should be recommended and why.

Summary & Review

Representing two categorical variables begins with the two-way table, which organizes raw counts into joint frequencies (interior cells) and marginal frequencies (row and column totals). From the table, we compute three types of proportions: joint relative frequencies (cell ÷ grand total), marginal relative frequencies (row or column total ÷ grand total), and conditional relative frequencies (cell ÷ row or column total). The conditional distributions are the key to detecting association.

Graphically, segmented bar charts and side-by-side bar charts are the primary tools for visualizing these relationships. Two variables are associated if the conditional distributions of one variable differ across the categories of the other; they are independent if the conditional distributions are approximately equal. Always remember: association does not imply causation, and beware of Simpson's Paradox, in which a lurking variable can reverse an apparent trend when data are aggregated.

Varsity Tutors • AP Statistics • Representing Two Categorical Variables