AP STATISTICS • INFERENCE FOR CATEGORICAL DATA: PROPORTIONS

Concluding a Test for a Population Proportion

Translate p-values and test statistics into well-justified statistical conclusions about a population proportion.

Historical Context & Motivation

The ability to draw rigorous conclusions from data did not emerge overnight; it developed through centuries of work by mathematicians and scientists grappling with uncertainty. The modern framework for hypothesis testing traces its roots to early probability theory and formalized statistical reasoning in the twentieth century. Understanding this history helps reveal why the conclusion step of a significance test—comparing a p-value to a significance level and writing a contextual interpretation—carries so much weight. Earlier statisticians recognized that simply computing a number was never enough; the real power of inference lies in translating that number into a defensible claim about the world.

1710
Arbuthnot's Significance Argument
John Arbuthnot analyzed London birth records and argued that the consistent excess of male births was too improbable to be mere chance—an early form of reasoning against a null hypothesis.
1900
Pearson's Chi-Square Test
Karl Pearson introduced the chi-square goodness-of-fit test, giving researchers a systematic way to compare observed proportions to expected proportions and quantify discrepancy.
1925
Fisher Formalizes P-Values
Ronald Fisher published 'Statistical Methods for Research Workers,' establishing the p-value as a continuous measure of evidence against the null hypothesis and popularizing the α = 0.05 threshold.
1933
Neyman–Pearson Framework
Jerzy Neyman and Egon Pearson introduced the concepts of Type I error, Type II error, and the formal reject/fail-to-reject decision rule that underpins modern hypothesis testing conclusions.
2019
ASA Statement on P-Values
The American Statistical Association issued supplementary guidance urging researchers to move beyond simple 'significant or not' conclusions—emphasizing context, effect size, and careful language in every inference.

This historical arc reveals a persistent theme: computing a test statistic or p-value is only part of the work. The conclusion step—stating whether the evidence is convincing, linking back to the original claim, and using proper statistical language—is what transforms arithmetic into inference. In this lesson we focus exclusively on that final, critical phase: how to correctly conclude a one-proportion z-test on the AP Statistics exam and in applied research.

Core Principles of Drawing a Conclusion

Drawing a conclusion from a one-proportion z-test requires you to connect a computed p-value (or z-statistic) back to the real-world question that motivated the test. The conclusion is not merely a yes-or-no verdict; it is a carefully worded statement that accounts for the hypotheses, the significance level, and the context of the data. The following foundational ideas govern this final step.

1

Compare p-value to α

If p-value ≤ α, reject H₀. If p-value > α, fail to reject H₀. The significance level α is the threshold set before data collection that defines 'convincing' evidence.
2

Never 'Accept' H₀

Failing to reject H₀ means the data did not provide sufficient evidence against it—not that H₀ is true. The correct language is 'fail to reject,' preserving the logical structure of proof by contradiction.
3

State the Conclusion in Context

A complete conclusion references the parameter, the population, and the direction of the alternative hypothesis in plain language. 'Reject H₀' alone earns no context credit on the AP exam.
4

Address Type I and Type II Error

A Type I error means rejecting H₀ when it is actually true. A Type II error means failing to reject H₀ when H_a is actually true. The conclusion you reach determines which error type is possible.
5

Link Evidence to the Alternative

When rejecting H₀, you conclude there is convincing evidence for H_a. When failing to reject, you conclude the evidence is not convincing enough—always in terms of the alternative claim.
KEY TAKEAWAY
Think of the significance level α as a pre-set alarm threshold on a smoke detector. If the smoke density (p-value) drops below the threshold, the alarm fires (reject H₀). If the smoke density stays above it, the alarm stays silent (fail to reject H₀). The alarm not going off does not prove there is no fire—it only means the detector did not sense enough smoke to trigger. Likewise, failing to reject H₀ never proves the null hypothesis is true; it only indicates insufficient evidence against it.

Visual Explanation: The Decision Diagram

The flowchart below maps the complete logic of concluding a one-proportion z-test. Starting from the computed p-value, it branches into the two possible decisions and shows the template language expected on the AP Statistics exam. Notice that every pathway ends with a contextual statement—this is non-negotiable for full credit.

The flowchart shows the two-branch decision process after computing a p-value. Each branch terminates with a conclusion template written in the language required by the AP Statistics exam, followed by the relevant error type.

The diagram highlights a crucial structural point: every conclusion has exactly two components. First, a formal decision (reject or fail to reject H₀) justified by the comparison of the p-value to α. Second, an interpretation in context that references the parameter, the population, and the direction of the alternative hypothesis. The AP scoring guidelines consistently award separate points for each component, so omitting either one will cost you credit.

Mathematical Framework for the Conclusion

Before you can write a conclusion, the test machinery must produce a test statistic and a p-value. Understanding these formulas is essential because your conclusion must logically follow from them. Below are the key equations and decision rules that lead to the final step.

ONE-PROPORTION Z-TEST STATISTIC
z = (p̂ − p₀) / √(p₀(1 − p₀) / n)
where is the sample proportion, p₀ is the hypothesized population proportion from H₀, and n is the sample size. The denominator is the standard error computed under the assumption that H₀ is true.
P-VALUE DEFINITION
p-value = P(observing a test statistic as extreme as, or more extreme than, z | H₀ is true)
For a one-sided test Hₐ: p > p₀, the p-value is the area to the right of z on the standard normal curve. For Hₐ: p < p₀, it is the area to the left. For a two-sided test Hₐ: p ≠ p₀, the p-value is twice the area in the more extreme tail.
DECISION RULE
If p-value ≤ α → Reject H₀; If p-value > α → Fail to reject H₀
The significance level α (commonly 0.05, 0.01, or 0.10) is specified before examining the data. The comparison is a strict inequality boundary: when the p-value exactly equals α, convention is to reject H₀.
📝 AP Exam Tip
On the AP Statistics exam, you must explicitly compare your p-value to α as part of your conclusion. Writing p-value = 0.023 < α = 0.05 before your reject/fail-to-reject statement ensures full credit for the linkage. Simply writing 'the p-value is small' without comparing it to a specific α is insufficient.

Detailed Breakdown: Writing the Conclusion Statement

The language of a statistical conclusion is precise and formulaic by design. The AP Statistics rubric consistently rewards students who hit specific wording checkpoints, and it penalizes common errors such as saying 'accept H₀' or failing to reference the context. This section dissects the anatomy of a correct conclusion statement and provides a visual reference for both reject and fail-to-reject scenarios.

The five required elements of a conclusion are shown as color-coded segments for the reject scenario, followed by complete example statements for both rejection and failure to reject. Note the common mistakes listed at the bottom.

Let us walk through each element shown in the diagram. Element 1 states the computed p-value as a number, providing transparency about the evidence. Element 2 is the comparison operator (< or >), which explicitly links the p-value to α. Element 3 names the significance level so the reader knows the threshold used. Element 4 gives the formal statistical decision: 'reject H₀' or 'fail to reject H₀.' Finally, Element 5 interprets the decision in the real-world context of the problem, referencing the parameter (the true population proportion), the population, and the direction stated in Hₐ. Omitting Element 5 is the single most common reason students lose points on the AP exam.

💬 Critical Vocabulary
Use 'convincing evidence' (or 'sufficient evidence') when rejecting H₀, and 'not convincing evidence' (or 'not sufficient evidence') when failing to reject. Never use the word 'prove' in a statistical conclusion—hypothesis tests assess probability, not certainty.

Worked Example: Full Hypothesis Test with Conclusion

A hospital administrator claims that fewer than 15% of patients discharged from the emergency department are readmitted within 30 days. A random sample of 200 discharged patients reveals that 22 were readmitted. At the α = 0.05 significance level, is there convincing evidence that the readmission rate is less than 15%?

One-Proportion z-Test for Hospital Readmission Rate
1
Step 1 — State the HypothesesDefine p as the true proportion of all emergency department patients at this hospital who are readmitted within 30 days. The null hypothesis is H₀: p = 0.15 and the alternative hypothesis is Hₐ: p < 0.15. This is a one-sided (left-tailed) test because the administrator's claim suggests the rate is below 15%.
2
Step 2 — Check ConditionsRandom: The sample is described as a random sample of 200 discharged patients. ✓ Independence: Assuming the hospital discharges far more than 10 × 200 = 2,000 patients, the 10% condition is met. ✓ Large Counts: np₀ = 200(0.15) = 30 ≥ 10 and n(1 − p₀) = 200(0.85) = 170 ≥ 10, so the normal approximation is appropriate. ✓
3
Step 3 — Compute the Test StatisticThe sample proportion is p̂ = 22/200 = 0.11. The standard error under H₀ is √(0.15 × 0.85 / 200) = √(0.1275 / 200) = √0.0006375 ≈ 0.02525. Therefore, z = (0.11 − 0.15) / 0.02525 = −0.04 / 0.02525 ≈ −1.584.
z ≈ −1.584
4
Step 4 — Find the P-ValueBecause this is a left-tailed test (Hₐ: p < 0.15), the p-value is P(Z ≤ −1.584). Using a standard normal table or calculator: p-value ≈ 0.0566.
p-value ≈ 0.0566
5
Step 5 — State the ConclusionBecause the p-value (0.0566) is greater than α = 0.05, we fail to reject H₀. There is not convincing evidence that the true proportion of emergency department patients at this hospital who are readmitted within 30 days is less than 0.15.
Fail to reject H₀ — not convincing evidence that p < 0.15
🔑 Why Step 5 Matters Most
Notice that the conclusion does not say 'we accept that p = 0.15' or 'the readmission rate is 15%.' It only says we lack convincing evidence that it is lower. The data are consistent with p = 0.15, but also consistent with many other values. This nuanced language is essential and reflects the Neyman–Pearson logic: the test is designed to control the probability of wrongly rejecting H₀, not to confirm it.

Type I Error, Type II Error, and Power

Every conclusion you draw in a hypothesis test carries the possibility of error. The type of error that is possible depends directly on the conclusion you reach. Understanding these errors—and being able to describe their consequences in context—is a required component of many AP free-response questions.

The four possible outcomes of a hypothesis test, showing which error type corresponds to each combination of truth and decision.
ScenarioDecisionError TypeConsequence in Context
H₀ is true, we reject H₀Reject H₀Type I ErrorWe conclude the readmission rate is below 15% when it actually is 15%—possibly leading the hospital to reduce follow-up resources prematurely.
H₀ is true, we fail to reject H₀Fail to reject H₀Correct decision ✓We correctly identify that there is not enough evidence the rate is below 15%.
Hₐ is true, we reject H₀Reject H₀Correct decision ✓We correctly detect that the readmission rate is below 15%.
Hₐ is true, we fail to reject H₀Fail to reject H₀Type II ErrorWe fail to detect that the rate is actually below 15%—missing an opportunity to recognize the hospital's improvement.
KEY TAKEAWAY
The relationship between Type I and Type II error is like a trade-off between a smoke detector's sensitivity and its false-alarm rate. If you lower α (make the detector less sensitive), you reduce the chance of a false alarm (Type I error) but increase the chance of missing a real fire (Type II error). Conversely, raising α increases sensitivity but also increases false alarms. Power—the probability of correctly rejecting H₀ when Hₐ is true—is the complement of the Type II error probability: Power = 1 − β. You can increase power by increasing the sample size, increasing α, or when the true parameter is farther from p₀.

Connection to Confidence Intervals and Two-Proportion Tests

The conclusion of a one-proportion z-test does not exist in isolation; it connects to the broader inferential framework you will encounter throughout AP Statistics. Two key connections deserve attention: the relationship between hypothesis tests and confidence intervals, and the extension to comparing two population proportions.

Comparing the one-proportion z-test conclusion to related inference procedures.
FeatureOne-Proportion z-TestOne-Proportion z-IntervalTwo-Proportion z-Test
Question answeredIs there convincing evidence that p differs from p₀?What is a plausible range for p?Is there convincing evidence that p₁ ≠ p₂?
Standard error usesp₀ (hypothesized value)p̂ (sample proportion)p̂ₓ (pooled proportion)
Outputz-statistic and p-valueInterval estimate (p̂ ± z*·SE)z-statistic and p-value
Conclusion formatReject or fail to reject H₀ in context"We are C% confident that p is between…"Reject or fail to reject H₀ about p₁ − p₂ in context

A powerful consistency check arises from the duality between hypothesis tests and confidence intervals: for a two-sided test at significance level α, rejecting H₀ is equivalent to observing that p₀ falls outside the corresponding (1 − α) × 100% confidence interval for p. If you construct a 95% confidence interval and it does not contain p₀, a two-sided test at α = 0.05 would reject H₀. This equivalence does not hold exactly for one-sided tests, but conceptually, the interval provides complementary information—estimating the parameter rather than testing a specific claim about it. As you progress to two-proportion inference, the conclusion template remains the same—compare p-value to α, make a decision, and interpret in context—but the parameter of interest shifts to the difference p₁ − p₂.

Practice Problems

1
A researcher conducts a one-proportion z-test and obtains a p-value of 0.03 with α = 0.05. Which of the following is the most appropriate conclusion?
2
A school official claims that 60% of students use the school bus. In a random sample of 150 students, 80 use the bus. The resulting one-proportion z-test yields z = −1.53 and a p-value of 0.063. At α = 0.05, what is the correct conclusion?
3
A consumer group suspects that more than 20% of a certain brand's cereal boxes are underweight. They test H₀: p = 0.20 versus Hₐ: p > 0.20 at α = 0.01. A random sample of 400 boxes yields p̂ = 0.25. The test statistic is z = 2.50 and the p-value is 0.0062. After stating the appropriate conclusion, identify which type of error could have been made and describe its consequence in context.
PROBLEM 4APPLIED
A pharmaceutical company claims that its new vaccine has an efficacy rate (proportion of vaccinated individuals who do not contract the disease) higher than 0.85. In a randomized clinical trial, 432 out of 500 vaccinated individuals did not contract the disease. The significance level is α = 0.05. (a) State the hypotheses. (b) Verify the conditions for inference. (c) Calculate the test statistic and p-value. (d) Write a complete conclusion in context, including an explicit comparison of the p-value to α. (e) Identify the type of error that could result from your conclusion and describe its consequence for public health.
PROBLEM 5CRITICAL THINKING
An environmental agency tests H₀: p = 0.10 versus Hₐ: p > 0.10, where p is the true proportion of water samples from a lake that exceed a safe lead concentration. A random sample of 300 water samples yields p̂ = 0.14 and a p-value of 0.031. (a) At α = 0.05, write a complete conclusion in context. (b) At α = 0.01, write a complete conclusion in context. (c) Explain why the same data can lead to different conclusions at different significance levels, and discuss how the agency should choose between α = 0.05 and α = 0.01 given the public health implications. (d) A colleague argues that because the sample proportion (0.14) is 'close' to 0.10, the result is practically unimportant even if statistically significant. Evaluate this argument, incorporating the concepts of statistical significance and practical significance.

Summary

Concluding a one-proportion z-test requires two integrated components. First, you compare the p-value to the pre-set significance level α and make a formal decision: reject H₀ if p-value ≤ α, or fail to reject H₀ if p-value > α. Second, you interpret the decision in the real-world context of the problem, referencing the population, the parameter, and the direction of Hₐ. Never say 'accept H₀' and never say 'prove'—the correct phrasing is 'there is (or is not) convincing evidence that…' followed by the alternative claim stated in plain language.

Identifying the potential error type is the final critical skill: when you reject, the risk is a Type I error (false rejection of a true H₀), and when you fail to reject, the risk is a Type II error (failure to detect a real departure from H₀). Describing these errors in context—explaining what real-world consequence would follow—demonstrates the deepest level of statistical reasoning and is essential for full credit on the AP Statistics exam.

Varsity Tutors • AP Statistics • Concluding a Test for a Population Proportion