What this quiz covers
This quiz focuses on Difference Of Two Means Test, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
A school district wants to know whether a new reading program changes mean reading-comprehension scores. A random sample of 40 students used the new program (Group N) and a separate random sample of 38 students used the old program (Group O). A two-sample t test for μN−μO was conducted with hypotheses H0:μN−μO=0 and Ha:μN−μO>0. The test produced a p-value of 0.018. Using α=0.05, what conclusion is appropriate?
AP Statistics Quiz
Practice Difference Of Two Means Test in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.
This quiz focuses on Difference Of Two Means Test, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.
A school district wants to know whether a new reading program changes mean reading-comprehension scores. A random sample of 40 students used the new program (Group N) and a separate random sample of 38 students used the old program (Group O). A two-sample t test for μN−μO was conducted with hypotheses H0:μN−μO=0 and Ha:μN−μO>0. The test produced a p-value of 0.018. Using α=0.05, what conclusion is appropriate?
Explanation: This question tests understanding of hypothesis test conclusions for a difference of two means. Since the p-value (0.018) is less than α (0.05), we reject the null hypothesis. The alternative hypothesis states μN - μO > 0, which means we're testing if the new program has a higher mean than the old program. When we reject H0, we have convincing evidence supporting the alternative hypothesis - that the new program increases the population mean score. Choice D incorrectly states we "prove" equality, but hypothesis tests never prove anything. Choice E overstates the conclusion by claiming the program works for "every student," when we can only make claims about population means.
A city compares mean response time (minutes) for two ambulance dispatch systems. Independent random samples of calls are taken from System 1 (S1) and System 2 (S2). A two-sample t test for μS1−μS2 is performed with H0:μS1−μS2=0 and Ha:μS1−μS2>0. The p-value is 0.11. Using α=0.10, what conclusion is appropriate?
Explanation: This question involves comparing ambulance response times with a right-tailed test. The p-value (0.11) is greater than α (0.10), so we fail to reject the null hypothesis. The alternative Ha: μS1 - μS2 > 0 tests whether System 1 has a larger (worse) mean response time than System 2. Since we fail to reject H0, we don't have convincing evidence that System 1 has a larger population mean response time. Choice D incorrectly interprets failing to reject as proof of equality - we simply lack evidence of a difference. Choice E makes an inappropriate causal claim. When p-value > α, we always fail to reject H0 and conclude there's insufficient evidence for the alternative hypothesis.
A school district wants to know whether a new online homework system changes mean weekly math quiz scores. A random sample of students used the new system (group N) and another random sample used the old system (group O). A two-sample t test was performed for H0:μN−μO=0 versus Ha:μN−μO>0. The test produced a p-value of 0.03. Using α=0.05, what conclusion is appropriate?
Explanation: This question tests your ability to interpret a two-sample t-test for the difference of means. Since the p-value (0.03) is less than the significance level α = 0.05, we reject the null hypothesis. The alternative hypothesis states μ_N - μ_O > 0, which means we're testing if the new system has a higher mean than the old system. By rejecting H₀, we conclude there is convincing evidence that the new system increases the mean quiz score. Choice D incorrectly claims we can conclude the population means are equal from sample data, and Choice E incorrectly makes a causal claim about individual students. When conducting hypothesis tests for difference of means, we compare the p-value to α and make conclusions about population means, not individual values.
A psychologist studies whether a mindfulness app affects mean stress score (higher = more stress). Participants were randomly assigned to use the app (group A) or a placebo app (group P). A two-sample t test was conducted for H0:μA−μP=0 versus Ha:μA−μP<0. The p-value was 0.018. At α=0.05, what conclusion is appropriate?
Explanation: This randomized experiment tests whether a mindfulness app lowers mean stress score, with H_a: μ_A - μ_P < 0. The p-value of 0.018 is less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the mindfulness app lowers mean stress score. Choice D incorrectly claims the app reduces stress for every individual, while Choice E wrongly states that we can determine the exact difference in population means. Even in randomized experiments, hypothesis tests provide evidence about population parameters (means), not guarantees about individual outcomes. The p-value tells us about statistical significance, not the size of the effect.
A school compares mean time (in minutes) to complete a standardized math test for students using a paper booklet versus an online version. Two independent random samples were taken (paper: n=40; online: n=38). A two-sample t test for μpaper−μonline used H0:μpaper−μonline=0 and Ha:μpaper−μonline<0. The p-value was 0.041. Using α=0.05, what conclusion is appropriate?
Explanation: This question tests the application of a two-sample t-test for the difference in mean completion times between paper and online math tests. The one-sided alternative hypothesizes that paper is faster, and since the p-value of 0.041 is less than α=0.05, we reject H0, supporting that the population mean time is lower for paper. A frequent distractor is choice B, which mistakenly says 'fail to reject' while claiming sufficient evidence, confusing the decision rule. When concluding two-mean tests, specify the direction if it's a one-sided test and evidence supports it, but don't generalize to causation like in choice E. Focus on population inferences, as sample differences alone (as in choice D) don't guarantee population differences without statistical significance.
An environmental scientist tests whether the mean nitrate concentration (mg/L) differs between two nearby lakes, Lake 1 and Lake 2. Independent random water samples were collected from each lake. A two-sample t test was performed for H0:μ1−μ2=0 versus Ha:μ1−μ2=0. The p-value was 0.049. Using α=0.05, what conclusion is appropriate?
Explanation: This two-sample t-test uses a two-tailed alternative hypothesis (μ₁ - μ₂ ≠ 0) to test for any difference in mean nitrate concentration between the lakes. The p-value of 0.049 is just barely less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the mean nitrate concentration differs between the two lakes. Choice C incorrectly specifies a direction (Lake 1 higher) when the two-tailed test doesn't indicate which lake has higher concentration. Choice E wrongly suggests that a p-value close to α makes the result inconclusive. In hypothesis testing, we use a clear decision rule: if p-value < α, we reject H₀, regardless of how close the values are.
A sports scientist compares mean resting heart rate for two populations: endurance athletes (A) and nonathletes (N). Independent random samples are taken, and a two-sample t test for μA−μN is performed with H0:μA−μN=0 and Ha:μA−μN<0. The p-value is 0.27. At the α=0.05 level, what conclusion is appropriate?
Explanation: This problem involves a left-tailed test for the difference between athlete and non-athlete mean heart rates. The p-value (0.27) is much larger than α (0.05), so we fail to reject the null hypothesis. The alternative hypothesis Ha: μA - μN < 0 tests whether athletes have a lower mean heart rate than non-athletes. Since we fail to reject H0, we do not have convincing evidence that athletes have a lower population mean resting heart rate. Choice D incorrectly claims that failing to reject H0 means the means are equal - we simply don't have evidence of a difference. Remember that failing to reject H0 never proves the null hypothesis is true; it only means we lack sufficient evidence against it.
A teacher compares mean quiz scores for students using two different study apps. Two independent random samples are taken: App A and App B. A two-sample t test for μA−μB is conducted with H0:μA−μB=0 and Ha:μA−μB=0. The p-value is 0.52. At α=0.05, what conclusion is appropriate?
Explanation: This problem presents a two-tailed test comparing study apps. The p-value (0.52) is much larger than α (0.05), so we fail to reject the null hypothesis. With a two-tailed alternative (Ha: μA - μB ≠ 0), failing to reject means we don't have convincing evidence of any difference between the population mean quiz scores. Choice A incorrectly suggests rejecting H0 when the large p-value clearly indicates we should fail to reject. Choice D wrongly claims this "proves" equality - hypothesis tests never prove the null hypothesis. A large p-value simply means our sample data are consistent with the null hypothesis of no difference between population means.
A school district compares mean math test scores for students taught with Method 1 versus Method 2. Two independent random samples of students were selected (Method 1: n=28; Method 2: n=30). A two-sample t test for μ1−μ2 was performed with H0:μ1=μ2 and Ha:μ1>μ2. The p-value was 0.004. Using α=0.01, what conclusion is appropriate?
Explanation: This question tests a one-tailed hypothesis where we're examining if Method 1 produces higher mean scores than Method 2 (H_a: μ₁ > μ₂). The p-value (0.004) is less than α = 0.01, so we reject the null hypothesis. This provides convincing evidence that Method 1 has a higher population mean score than Method 2. Choice C incorrectly claims the effect applies to "every student," which is an overstatement—hypothesis tests make conclusions about population means, not individual outcomes. Choice D reverses the direction of the conclusion. When conducting hypothesis tests, we make inferences about population parameters based on sample data, not claims about every individual in the population.
A public health researcher compares mean systolic blood pressure for adults who exercise regularly versus those who do not. Two independent random samples were taken (Exercise: n=50; No exercise: n=48). A two-sample t test for μex−μno was conducted with H0:μex=μno and Ha:μex<μno. The p-value was 0.031. Using α=0.05, what conclusion is appropriate?
Explanation: This question tests understanding of a one-tailed test where we're examining if exercisers have lower mean systolic blood pressure (H_a: μ_ex < μ_no). The p-value (0.031) is less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the population mean systolic blood pressure is lower for adults who exercise regularly than for those who do not. Choice C incorrectly claims causation for "every adult," which overstates the conclusion. Choice E misinterprets what failing to reject would mean. When we reject H₀ in a one-tailed test with "less than" alternative, we conclude there's evidence supporting the specific directional claim in the alternative hypothesis.
A researcher compares mean daily screen time (hours) for teenagers in Urban areas versus Rural areas. Independent random samples were taken (Urban: n=45; Rural: n=43). A two-sample t test for μU−μR used hypotheses H0:μU=μR and Ha:μU=μR. The p-value was 0.051. Using α=0.05, what conclusion is appropriate?
Explanation: This problem involves a two-tailed test comparing mean screen time between Urban and Rural teenagers. The p-value (0.051) is just slightly greater than α = 0.05, so we fail to reject the null hypothesis. This means there is not convincing evidence at the 0.05 level that population mean screen time differs between Urban and Rural teenagers. Choice D incorrectly concludes that the population means are equal—failing to reject H₀ doesn't prove equality. Choice C makes an inappropriate causal claim. When the p-value is very close to α but slightly exceeds it, we still fail to reject H₀, though the evidence is nearly significant at the chosen level.
An engineer compares mean battery life (hours) for two brands, Brand X and Brand Y. Independent random samples were tested (Brand X: n=18; Brand Y: n=20). A two-sample t test for μX−μY used H0:μX=μY and Ha:μX=μY. The p-value was 0.62. At α=0.05, what conclusion is appropriate?
Explanation: This problem involves a two-tailed test comparing mean battery life between two brands. With a p-value of 0.62, which greatly exceeds α = 0.05, we fail to reject the null hypothesis. This means there is not convincing evidence of a difference in population mean battery life between Brand X and Brand Y. Choice C incorrectly states that the brands have "exactly the same mean"—failing to reject H₀ doesn't prove equality, it only indicates insufficient evidence for a difference. Choice E makes an inappropriate causal claim. In hypothesis testing, a large p-value suggests the observed difference could easily occur by chance if the null hypothesis were true, so we maintain our assumption of no difference.
A company tests whether a new training program reduces the mean time (in minutes) to complete a standard task. A random sample of employees used the new program (group T) and a separate random sample used the old program (group C). A two-sample t test was performed for H0:μT−μC=0 versus Ha:μT−μC<0. The p-value was 0.08. Using α=0.10, what conclusion is appropriate?
Explanation: In this two-sample t-test, we're testing whether the new training program reduces mean completion time, with H_a: μ_T - μ_C < 0 (new minus old is negative). The p-value of 0.08 is less than the significance level α = 0.10, so we reject the null hypothesis. This provides convincing evidence that the new program reduces the mean completion time. Choice D incorrectly makes a claim about every individual employee rather than the population mean. Choice E wrongly states that failing to reject H₀ would prove no difference exists. When interpreting hypothesis tests, focus on whether the p-value is less than α, and remember that conclusions apply to population parameters, not individual observations.
A bookstore owner compares mean amount spent per customer on weekdays versus weekends. Independent random samples were taken (weekday: n=25; weekend: n=27). A two-sample t test for μweekend−μweekday used H0:μweekend−μweekday=0 and Ha:μweekend−μweekday>0. The p-value was 0.073. At α=0.05, what conclusion is appropriate?
Explanation: This question tests a two-sample t-test for comparing mean spending between weekdays and weekends. The one-sided alternative posits higher weekend spending, but p=0.073 > α=0.05 leads to failing to reject H0, with insufficient evidence for the claim. Choice D distracts by claiming fail to reject proves equal means, but it doesn't—it just lacks evidence of difference. Two-mean test conclusions should carefully state what the evidence supports regarding population means, avoiding causal or definitive equality statements like in choice E. Always compare p to α explicitly in your reasoning.
A city compares mean commute time (minutes) for residents who take the bus versus residents who drive. Two independent random samples were selected (bus: n=55; drive: n=50). A two-sample t test for μbus−μdrive used hypotheses H0:μbus−μdrive=0 and Ha:μbus−μdrive=0. The p-value was 0.004. Using α=0.01, what conclusion is appropriate?
Explanation: This question examines a two-sample t-test for differences in mean commute times between bus riders and drivers. The two-sided alternative allows for any difference, and with p=0.004 less than α=0.01, we reject H0, providing evidence of a population mean difference. Choice D is a distractor, incorrectly assuming different sample means automatically mean different population means without considering significance. For two-mean tests, conclusions should avoid causal claims like in choice E and focus on evidence for or against the null in the population context. Note that even with rejection, we don't specify direction unless the test is one-sided or further analysis is done.
An engineer compares mean battery life (hours) for two brands of rechargeable batteries. Independent random samples were tested (Brand X: n=15; Brand Y: n=17). A two-sample t test for μX−μY was conducted with H0:μX−μY=0 and Ha:μX−μY>0. The p-value was 0.11. At α=0.05, what conclusion is appropriate?
Explanation: This question involves interpreting a two-sample t-test for comparing mean battery life between two brands. With a one-sided alternative for Brand X being greater and p-value 0.11 exceeding α=0.05, we fail to reject H0, meaning not enough evidence that X has longer life. Choice D distracts by saying fail to reject proves equal means, but it only indicates insufficient evidence against equality. In two-mean test conclusions, remember that failing to reject doesn't affirm the null or imply causation, as choice E wrongly suggests for Brand Y. Always reference the population means and the specific alternative hypothesis in your statement.
A nutrition researcher compares mean systolic blood pressure for adults who follow Diet A versus Diet B. Two independent random samples were taken (Diet A: n=22; Diet B: n=24). A two-sample t test for μA−μB was performed with H0:μA−μB=0 and Ha:μA−μB=0. The p-value was 0.62. At the 0.05 significance level, what conclusion is appropriate?
Explanation: This question evaluates understanding of a two-sample t-test for comparing means of systolic blood pressure between two diets. The two-sided alternative tests for any difference, and with a p-value of 0.62 greater than α=0.05, we fail to reject the null hypothesis, indicating insufficient evidence of a population mean difference. Choice D is a distractor because it claims failing to reject proves exact equality in population means, but it only means we lack evidence to say they differ. In conclusions for two-mean tests, emphasize that 'fail to reject' doesn't confirm the null—there might still be a difference undetected by the sample. Avoid implying causation, as in choice E, since the test doesn't establish cause-and-effect relationships. Always tie the conclusion back to the population parameters, not just the samples.
A sports scientist investigates whether mean vertical jump height differs between athletes who follow Program 1 and those who follow Program 2. Two independent random samples were taken (Program 1: n=12; Program 2: n=14). A two-sample t test for μ1−μ2 was conducted with H0:μ1−μ2=0 and Ha:μ1−μ2=0. The p-value was 0.049. Using α=0.05, what conclusion is appropriate?
Explanation: This question assesses interpreting a two-sample t-test for mean jump height differences between two programs. The two-sided test yields p=0.049 just under α=0.05, so we reject H0, indicating sufficient evidence of a population mean difference. A common distractor is choice D, which confuses sample mean differences with the hypothesis test outcome. In concluding two-mean tests, emphasize population-level inferences and avoid absolutes like 'for all athletes' in choice E, as results are probabilistic. Rejection supports the alternative but doesn't prove causation or direction without additional context.
A gardener compares mean plant height after 8 weeks for plants given Fertilizer 1 versus Fertilizer 2. Two independent random samples were used (Fertilizer 1: n=20; Fertilizer 2: n=20). A two-sample t test for μ1−μ2 was performed with H0:μ1−μ2=0 and Ha:μ1−μ2=0. The p-value was 0.032. Using α=0.05, what conclusion is appropriate?
Explanation: This question tests interpreting a two-sample t-test for mean plant heights with different fertilizers. The two-sided alternative and p=0.032 < α=0.05 lead to rejecting H0, showing evidence of a population mean difference. Choice D is a distractor, wrongly equating sample differences to automatic population differences without significance. In two-mean conclusions, avoid causal language like choice E and focus on evidence for the alternative in the population. Rejection indicates a difference but not its direction in a two-sided test unless specified.
A nutritionist compares mean sodium intake (mg/day) for adults who follow Diet A versus Diet B. Independent random samples were taken from each diet group. A two-sample t test was conducted for H0:μA−μB=0 versus Ha:μA−μB=0. The p-value was 0.41. At the 0.05 significance level, what conclusion is appropriate?
Explanation: This problem involves a two-sample t-test with a two-tailed alternative hypothesis (μ_A - μ_B ≠ 0). The p-value of 0.41 is much larger than the significance level of 0.05, so we fail to reject the null hypothesis. This means there is not convincing evidence that mean sodium intake differs between Diet A and Diet B. Choice D incorrectly suggests that similar sample means guarantee equal population means, while Choice E wrongly claims that failing to reject H₀ proves the means are exactly equal. When we fail to reject H₀, we simply lack sufficient evidence to conclude the means differ; we cannot prove they are equal. Remember that hypothesis tests can only provide evidence against the null hypothesis, never proof that it's true.