What this quiz covers
This quiz focuses on Statistical Inference, giving you a quick way to practice the rules, question types, and explanations that matter most for USMLE Step 1.
Researchers investigate the effect of a new statin on LDL cholesterol levels. In a clinical trial, the mean difference in LDL reduction between the statin group and the placebo group was 25 mg/dL. The 95% confidence interval for this mean difference was calculated to be (15 mg/dL, 35 mg/dL).
Which of the following is the most accurate interpretation of this 95% confidence interval?
USMLE Step 1 Quiz
Practice Statistical Inference in USMLE Step 1 with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.
This quiz focuses on Statistical Inference, giving you a quick way to practice the rules, question types, and explanations that matter most for USMLE Step 1.
Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.
Researchers investigate the effect of a new statin on LDL cholesterol levels. In a clinical trial, the mean difference in LDL reduction between the statin group and the placebo group was 25 mg/dL. The 95% confidence interval for this mean difference was calculated to be (15 mg/dL, 35 mg/dL).
Which of the following is the most accurate interpretation of this 95% confidence interval?
Explanation: A 95% confidence interval provides a range of plausible values for the true population parameter. The correct frequentist interpretation is that if the same study were conducted 100 times, 95 of the resulting confidence intervals would be expected to contain the true mean difference in the population.
A study is designed to determine if a new educational intervention improves medical student scores on a standardized exam. The researchers set the significance level (α) to 0.05. After the intervention, they find that the intervention group scores significantly higher than the control group (p=0.04) and conclude the intervention is effective. However, in reality, the intervention has no true effect on exam scores.
In this scenario, the researchers have made which of the following types of error?
Explanation: A Type I error occurs when the null hypothesis is incorrectly rejected. The null hypothesis states there is no difference between the groups. Here, the researchers concluded there was a difference (rejected the null) when in fact no true difference existed. This is the definition of a Type I error (false positive). The probability of making a Type I error is denoted by α.
A pharmaceutical company develops a new medication intended to prevent migraines. A clinical trial is conducted, but the results show no statistically significant difference in migraine frequency between the medication group and the placebo group (p = 0.15). The company decides not to pursue further development. Later, several larger studies confirm that the medication is, in fact, modestly effective.
The initial study's failure to detect a real effect is an example of which of the following?
Explanation: A Type II error occurs when one fails to reject a null hypothesis that is actually false. In this case, the initial study failed to find a significant effect (did not reject the null hypothesis of no difference) when a true effect existed. This is the definition of a Type II error (false negative). The probability of making a Type II error is denoted by β.
A public health study reports the annual income of residents in a community with a large academic medical center. The data show that a few highly paid surgeons earn substantially more than the majority of residents, who have modest incomes. The distribution of income is highly skewed.
Which of the following measures of central tendency would be the most appropriate to describe the typical income in this community?
Explanation: In a skewed distribution, the mean is heavily influenced by extreme values (outliers). The median represents the 50th percentile and is resistant to outliers. Therefore, for skewed data such as income, the median provides a more accurate representation of the central or 'typical' value than the mean.
A research team is planning a study to compare a new drug for diabetes with a standard treatment. They want to ensure their study has a high probability of detecting a clinically meaningful difference in HbA1c levels if one truly exists. They aim for a study power of 80%.
Which of the following is the most effective way for the researchers to increase the power of their study?
Explanation: Power is the ability of a study to detect a true effect (1 - β). The power of a study is primarily influenced by the sample size, effect size, and alpha level. Increasing the sample size is the most common and effective method to increase statistical power, as it reduces the standard error and makes it easier to detect a true difference between groups.
A study is published on the weights of newborns in a specific hospital. The mean weight is 3400 grams, with a standard deviation of 400 grams. A second publication reports on the mean birth weight from 100 different hospitals, giving a mean of 3400 grams and a standard error of the mean of 40 grams.
Which of the following best describes the standard error of the mean (SEM)?
Explanation: The standard deviation (SD) measures the variability or spread of individual data points within a single sample. The standard error of the mean (SEM) estimates the variability of the means of multiple samples taken from the same population. It quantifies how precisely the sample mean estimates the true population mean and is calculated as SD / √n.
A survey asks physicians to report the number of hours they sleep per night. The results are plotted, and the distribution is found to have a tail extending to the left. The peak of the distribution is at 8 hours, but a number of physicians report sleeping only 4 or 5 hours, pulling the average down.
Which of the following best describes the relationship between the measures of central tendency for this distribution?
Explanation: This describes a negatively skewed (left-skewed) distribution. In such a distribution, the outliers are on the lower end. The mode is the most frequent value (the peak), which is 8 hours. The median is less affected by the low-value outliers than the mean. The mean is pulled downward by the low values. Therefore, the relationship is Mean < Median < Mode.
Two different studies evaluate the same new surgical procedure. Study A has 50 participants, while Study B has 5000 participants. Both studies find the same point estimate for the reduction in recovery time and are free of bias. Assume all other factors are equal.
How would the 95% confidence interval (CI) for the mean reduction in recovery time in Study B compare to that of Study A?
Explanation: The width of a confidence interval is inversely related to the square root of the sample size. A larger sample size (like in Study B) leads to a smaller standard error of the mean, resulting in a more precise estimate of the true population parameter. This increased precision is reflected by a narrower confidence interval.
A clinical trial compares a new cancer therapy to a standard therapy. The primary endpoint is 5-year survival. The results show a 5-year survival of 45% with the new therapy and 42% with the standard therapy. The p-value for this difference is 0.04. The researchers claim the new therapy is superior.
Which of the following statements represents the most critical consideration when interpreting this result?
Explanation: While the result is statistically significant (p < 0.05), the absolute difference in survival is only 3%. A critical step in interpreting research is to evaluate whether a statistically significant finding is also clinically meaningful. A small, clinically unimportant difference can become statistically significant if the sample size is very large. Clinicians must decide if a 3% survival benefit justifies the potential costs, side effects, and risks of the new therapy.
Before starting a randomized controlled trial for a new drug, the investigators must specify the probability of making a Type I error that they are willing to accept. This threshold is used to determine statistical significance at the end of the study.
This pre-specified probability threshold is known as which of the following?
Explanation: The alpha (α) level, or significance level, is the pre-specified probability of committing a Type I error. It is the threshold below which a p-value is considered statistically significant. By convention, α is typically set to 0.05, meaning the researchers accept up to a 5% chance of incorrectly rejecting a true null hypothesis.
Researchers are planning a cohort study to investigate the effect of a new lifestyle intervention on the incidence of type 2 diabetes. They anticipate that the effect of the intervention will be small. They want to ensure their study has adequate power to detect this small effect.
Besides increasing the sample size, which of the following would also increase the power of the study?
Explanation: Power (1-β) is the probability of correctly rejecting a false null hypothesis. There is a trade-off between Type I (α) and Type II (β) errors. By increasing the alpha level (e.g., from 0.05 to 0.10), the threshold for significance becomes less strict, making it easier to reject the null hypothesis. This reduces the probability of a Type II error (β) and therefore increases the power of the study, at the cost of increasing the risk of a Type I error.
A clinical laboratory establishes a reference range for serum potassium. They measure the level in thousands of healthy volunteers and find that the values are normally distributed. The reference range is defined as the central 95% of these values.
This reference range corresponds to which of the following statistical intervals?
Explanation: For a normally distributed variable, approximately 95% of all individual values lie within 2 standard deviations (more precisely, 1.96 SDs) of the population mean. This is the standard definition used for creating reference ranges for many laboratory tests. A 95% confidence interval describes the range for the population mean, not the range for individual values.
A large pharmaceutical company conducts a well-designed, randomized trial to test a new cholesterol-lowering drug. The results show a reduction in LDL cholesterol that is not statistically significant (p = 0.25). The company shelves the drug. A junior researcher argues that because the trial had a relatively small sample size, a clinically important effect might have been missed.
The researcher is suggesting that the study may have resulted in which type of error due to insufficient power?
Explanation: The study failed to reject the null hypothesis (p > 0.05). The researcher's concern is that a true effect exists but was not detected. This scenario—failing to detect a real effect—is a Type II error. Studies with small sample sizes often lack sufficient statistical power to detect small or moderate effects, increasing the risk of a Type II error.
A new rapid screening test for a certain viral infection is developed. In a trial, the test fails to identify the infection in a small number of patients who are later confirmed to have the disease by a gold-standard PCR test. The company is concerned about the consequences of these false negatives.
The failure of the test to detect a disease when it is truly present corresponds to which statistical concept?
Explanation: In the context of diagnostic testing, the null hypothesis is that the patient does not have the disease. A false negative occurs when the test result is negative, but the disease is present. This is analogous to a Type II error: failing to reject the null hypothesis (of no disease) when it is false (disease is present). Low sensitivity, not specificity, is the test characteristic associated with high rates of false negatives.
A randomized controlled trial is conducted to evaluate the efficacy of a new antihypertensive drug, Drug X, compared to a placebo. After 12 weeks, the mean reduction in systolic blood pressure (SBP) was 10 mmHg in the Drug X group and 2 mmHg in the placebo group. The difference in mean SBP reduction between the two groups was 8 mmHg. A statistical analysis yields a p-value of 0.03 for this difference.
Based on this p-value, which of the following is the most appropriate conclusion?
Explanation: The p-value is the probability of observing a result at least as extreme as the one obtained, assuming the null hypothesis is true. In this case, the null hypothesis is that there is no difference in SBP reduction between Drug X and placebo. Therefore, a p-value of 0.03 means there is a 3% chance of seeing a difference of 8 mmHg or greater if the drug truly has no effect.
The fasting glucose levels of a healthy adult population are known to be approximately normally distributed with a mean of 90 mg/dL and a standard deviation of 8 mg/dL.
Based on this information, approximately what percentage of this population would be expected to have a fasting glucose level between 74 mg/dL and 106 mg/dL?
Explanation: For a normal distribution, approximately 68% of the data falls within 1 standard deviation (SD) of the mean, 95% falls within 2 SDs, and 99.7% falls within 3 SDs. The range from 74 mg/dL to 106 mg/dL represents the mean (90) plus or minus 16 mg/dL. Since the SD is 8 mg/dL, this range corresponds to the mean ± 2 SDs (90 ± 2*8). Therefore, approximately 95% of the population falls within this range.
A study measures the incubation period for a viral illness in a group of 1,000 infected individuals. The mean incubation period is 10 days, the median is 8 days, and the mode is 7 days. The distribution is asymmetrical.
Based on these measures of central tendency, what is the shape of this distribution?
Explanation: In a positively skewed (right-skewed) distribution, there is a long tail of high values. These high values pull the mean to the right of the median. The mode is the most frequent value and is typically at the peak of the distribution, to the left of the median. The relationship Mode < Median < Mean (7 < 8 < 10) is characteristic of a positively skewed distribution.
A researcher is studying serum cholesterol levels in a population of healthy adults. The data is collected and analyzed from 1000 participants. The distribution of cholesterol levels is found to be symmetrical and bell-shaped.
For this distribution, which of the following relationships between the mean, median, and mode is expected?
Explanation: A symmetrical, bell-shaped distribution is characteristic of a normal distribution. In a perfect normal distribution, the data is symmetrically distributed around the center. As a result, the mean (the arithmetic average), the median (the middle value), and the mode (the most frequent value) are all equal and located at the center of the distribution.
A case-control study investigated the association between daily consumption of a specific artificial sweetener and the risk of bladder cancer. The study found an odds ratio of 1.5. The 95% confidence interval for the odds ratio was (0.9, 2.5). The p-value was 0.08.
Which of the following is the best interpretation of these findings?
Explanation: For an odds ratio or relative risk, the null value is 1.0, which indicates no association between the exposure and the outcome. Since the 95% confidence interval (0.9, 2.5) includes 1.0, the result is not statistically significant at the α = 0.05 level. This is consistent with the p-value of 0.08, which is greater than 0.05.
A study reports a mean serum sodium level of 140 mEq/L in a sample of 100 patients. The standard deviation (SD) is 5 mEq/L. The researchers also report the 95% confidence interval for the true mean sodium level in the population from which the sample was drawn.
The calculation of this confidence interval is based on the point estimate (the sample mean) and which of the following measures of variability?
Explanation: A confidence interval for a mean is calculated as: Sample Mean ± (Critical Value * SEM). The standard error of the mean (SEM) is used because the goal is to estimate the precision of the sample mean as an estimate of the true population mean. The SEM (calculated as SD/√n) quantifies this uncertainty. The SD measures variability of individual data points, not the precision of the mean.