USMLE STEP 1 • BIOSTATISTICS AND EPIDEMIOLOGY

Risk Measures And Screening

Quantifying disease risk and evaluating screening tests are fundamental skills for evidence-based clinical decision-making.

Historical Context & Motivation

The ability to quantify disease risk and evaluate the accuracy of diagnostic tests did not emerge from a single eureka moment; rather, it evolved over centuries as physicians moved from anecdotal reasoning toward evidence-based medicine. Early epidemiologists recognized that simply counting cases was insufficient—what mattered was relating those cases to the population at risk, the time period of observation, and the exposures that might explain disease occurrence. This conceptual shift gave rise to measures such as relative risk, odds ratio, and attributable risk, each answering a slightly different clinical question about how exposure and disease are linked.

1854
John Snow & Cholera
John Snow mapped cholera cases in London's Broad Street outbreak, comparing attack rates between water sources—an early application of risk measurement in epidemiology.
1951
Doll & Hill Cohort Study
Richard Doll and A. Bradford Hill launched a landmark cohort study of British physicians, formalizing relative risk as a measure linking smoking to lung cancer.
1966
Wilson & Jungner Screening Criteria
The WHO published criteria for evaluating screening programs, establishing the importance of sensitivity, specificity, and predictive values in public health decision-making.
1975
Sackett's Bias Framework
David Sackett cataloged systematic biases in clinical studies, including lead-time bias and length-time bias, which are critical for interpreting screening outcomes.
1996
Evidence-Based Medicine Movement
The EBM working group formalized quantitative risk assessment—including number needed to treat—as integral to clinical decision-making at the bedside.

These historical milestones converge on a central question that the USMLE expects you to answer fluently: given a clinical or research scenario, which risk measure best quantifies the relationship between exposure and disease, and how do you evaluate whether a screening test is worth implementing? Mastering these tools is not merely an exercise in arithmetic—it directly informs whether a clinician recommends a mammogram, prescribes a statin, or counsels a patient about occupational hazards.

Core Principles & Definitions

Before diving into formulas, it is essential to understand the conceptual architecture that underpins all risk measures and screening metrics. Every risk measure begins with a clear definition of who is exposed versus unexposed, and who develops the outcome of interest. The classic 2 × 2 contingency table organizes these counts into four cells, from which virtually every epidemiological measure can be derived. Screening tests add a second dimension: the relationship between a test result and the true disease state, captured by sensitivity, specificity, and predictive values.

1

Relative Risk (RR)

The ratio of disease incidence in the exposed group to that in the unexposed group. Used in cohort studies and randomized controlled trials to quantify the strength of association.
2

Odds Ratio (OR)

The ratio of the odds of exposure among cases to the odds of exposure among controls. The primary measure of association in case-control studies, and it approximates RR when disease prevalence is low.
3

Attributable Risk (AR)

The absolute difference in disease incidence between exposed and unexposed groups. It tells clinicians how much additional risk is due to the exposure itself.
4

Number Needed to Treat / Harm (NNT / NNH)

The reciprocal of the absolute risk reduction (or increase). NNT tells you how many patients you must treat to prevent one additional adverse outcome—a clinically intuitive metric.
5

Sensitivity & Specificity

Sensitivity is the probability of a positive test given disease (SnNOut: high sensitivity rules out). Specificity is the probability of a negative test given no disease (SpPIn: high specificity rules in).
KEY TAKEAWAY
Think of relative risk as a magnifying glass—it tells you how much more likely an exposed person is to develop disease compared to an unexposed person. Attributable risk, on the other hand, is like a ruler—it tells you the exact extra amount of disease due to that exposure. A magnifying glass (RR = 10) looks dramatic, but if the baseline risk is tiny, the ruler (AR) may show the excess is clinically negligible. Always pair relative and absolute measures when making clinical decisions.

Visual Explanation — The 2 × 2 Table

The 2 × 2 contingency table is the single most important organizational tool in epidemiological risk assessment. Every risk measure—relative risk, odds ratio, attributable risk, sensitivity, specificity, and predictive values—derives directly from four cells labeled a, b, c, and d. The following diagram illustrates how these cells are arranged and how each measure maps onto them.

The 2 × 2 table organizes subjects by exposure (rows) and disease status (columns). Cells a, b, c, and d are the building blocks for every risk measure and screening metric shown below the dashed line.

Notice that for risk measures (relative risk, attributable risk), the rows represent exposure status—exposed versus unexposed—while the columns represent disease outcome. For screening metrics (sensitivity, specificity, PPV, NPV), the same table can be reframed: rows become test result (positive vs. negative) and columns become true disease state. This dual interpretation is why the 2 × 2 table appears so frequently on the USMLE—it is the universal scaffold upon which nearly every biostatistics question is built.

Mathematical Framework

Each risk measure answers a distinct clinical question. Below are the formal definitions along with variable explanations. In every formula, a = exposed with disease, b = exposed without disease, c = unexposed with disease, and d = unexposed without disease.

RELATIVE RISK (RR)
RR = [a / (a + b)] ÷ [c / (c + d)]
The numerator is the incidence in the exposed group; the denominator is the incidence in the unexposed group. RR = 1 means no association; RR > 1 suggests increased risk with exposure; RR < 1 suggests a protective effect. Used in cohort studies and RCTs.
ODDS RATIO (OR)
OR = (a × d) ÷ (b × c)
The odds ratio compares the odds of exposure among cases to the odds of exposure among controls. It is the measure of association for case-control studies. When disease prevalence is < 10%, the OR closely approximates the RR (the rare disease assumption).
ATTRIBUTABLE RISK (AR) / RISK DIFFERENCE
AR = [a / (a + b)] − [c / (c + d)]
AR quantifies the absolute excess risk due to exposure. It is clinically meaningful because it tells you how many additional cases per unit population are attributable to the exposure. The Number Needed to Treat (NNT) = 1 / AR (or 1 / absolute risk reduction).
SENSITIVITY & SPECIFICITY
Sensitivity = a / (a + c) | Specificity = d / (b + d)
Sensitivity (true positive rate) answers: of all people with disease, what proportion tested positive? Specificity (true negative rate) answers: of all people without disease, what proportion tested negative? These are intrinsic test properties and do not change with disease prevalence.
💡 HIGH-YIELD USMLE TIP
Remember the mnemonics: SnNOut — a highly Sensitive test, when Negative, rules Out disease. SpPIn — a highly Specific test, when Positive, rules In disease.

Screening Metrics & the Effect of Prevalence

While sensitivity and specificity are fixed properties of a test, the positive predictive value (PPV) and negative predictive value (NPV) depend critically on the prevalence of disease in the population being tested. PPV = a / (a + b) represents the probability that a person with a positive test actually has the disease, while NPV = d / (c + d) represents the probability that a person with a negative test is truly disease-free. As prevalence increases, PPV rises and NPV falls; as prevalence decreases, PPV drops (more false positives relative to true positives) and NPV increases. This explains why screening for rare diseases in the general population produces many false alarms.

At low prevalence (e.g., 1%), even a test with 90% sensitivity and 95% specificity yields a PPV of only ~15%. As prevalence rises to 40%, PPV climbs to ~92%. Meanwhile, NPV remains very high at low prevalence but declines as disease becomes more common.

This prevalence-dependence has direct clinical implications. A screening test for a rare condition (e.g., phenylketonuria in newborns) must have extremely high specificity to keep the false-positive rate manageable. Conversely, when prevalence is high—such as screening for HIV in high-risk populations—even tests with modest specificity can achieve respectable PPV. The USMLE frequently tests this concept by asking what happens to PPV or NPV when the same test is applied to populations with different baseline disease prevalences.

📌 PREVALENCE RULE OF THUMB
↑ Prevalence → ↑ PPV, ↓ NPV. ↓ Prevalence → ↓ PPV, ↑ NPV. Sensitivity and specificity remain unchanged because they are intrinsic to the test, not the population.

Worked Example

A cohort study follows 2,000 factory workers for 10 years: 800 are exposed to asbestos and 1,200 are unexposed. Among the exposed group, 48 develop mesothelioma. Among the unexposed group, 12 develop mesothelioma. Calculate the relative risk, attributable risk, and number needed to harm (NNH).

Asbestos Exposure & Mesothelioma Risk
1
Step 1 — Construct the 2 × 2 TableExposed with disease (a) = 48. Exposed without disease (b) = 800 − 48 = 752. Unexposed with disease (c) = 12. Unexposed without disease (d) = 1,200 − 12 = 1,188.
a = 48, b = 752, c = 12, d = 1,188
2
Step 2 — Calculate Incidence in Each GroupIncidence in exposed = a / (a + b) = 48 / 800 = 0.06 (6%). Incidence in unexposed = c / (c + d) = 12 / 1,200 = 0.01 (1%).
Incidenceexposed = 6%, Incidenceunexposed = 1%
3
Step 3 — Calculate Relative Risk (RR)RR = 0.06 / 0.01 = 6.0. This means exposed workers are 6 times more likely to develop mesothelioma compared to unexposed workers.
RR = 6.0
4
Step 4 — Calculate Attributable Risk (AR)AR = 0.06 − 0.01 = 0.05 (5%). For every 100 exposed workers, 5 additional cases of mesothelioma can be attributed to asbestos exposure beyond the background rate.
AR = 5% (0.05)
5
Step 5 — Calculate Number Needed to Harm (NNH)NNH = 1 / AR = 1 / 0.05 = 20. This means that for every 20 workers exposed to asbestos, one additional case of mesothelioma is expected beyond what would occur without exposure.
NNH = 20

Strengths, Limitations & Comparisons

Comparison of key risk measures and screening metrics
MeasureStrengthsLimitations
Relative RiskIntuitive ratio interpretation; directly calculated from cohort data; applicable to RCTs.Cannot be calculated from case-control studies; may overstate importance when baseline risk is very low.
Odds RatioThe only measure of association available in case-control studies; approximates RR when disease is rare.Less intuitive than RR; overestimates RR when prevalence is high; does not directly give incidence.
Attributable RiskProvides the absolute excess risk; clinically actionable; basis for NNT/NNH calculations.Requires incidence data (cohort/RCT only); does not convey fold-change in risk.
SensitivityIdeal for screening; a negative result with high sensitivity essentially rules out disease (SnNOut).High sensitivity alone does not mean the test is useful—specificity must also be considered; does not change with prevalence but PPV does.
SpecificityIdeal for confirmation; a positive result with high specificity rules in disease (SpPIn).A very specific test may miss cases (low sensitivity); increasing specificity generally decreases sensitivity.
🩺 CLINICAL INTEGRATION
Think of the diagnostic process like a two-stage security check at an airport. The first gate (screening test with high sensitivity) casts a wide net—it flags everyone who might be a threat, knowing some innocent passengers will be stopped. The second gate (confirmatory test with high specificity) narrows it down—only the true threats are detained. In medicine, you screen broadly (e.g., ELISA for HIV), then confirm narrowly (e.g., Western blot). This two-step approach maximizes both clinical safety and resource efficiency.

Connection to Advanced Concepts — Likelihood Ratios & ROC Curves

While sensitivity, specificity, and predictive values form the foundation of screening evaluation, more advanced tools allow clinicians to refine diagnostic reasoning. The likelihood ratio (LR) combines sensitivity and specificity into a single metric that is independent of prevalence. The positive likelihood ratio (LR+) = Sensitivity / (1 − Specificity), while the negative likelihood ratio (LR−) = (1 − Sensitivity) / Specificity. A high LR+ (>10) dramatically increases the post-test probability of disease, while a low LR− (<0.1) dramatically decreases it.

Basic versus advanced risk and screening frameworks
ConceptBasic FrameworkAdvanced Extension
Test AccuracySensitivity and specificity reported as separate valuesROC curve plots sensitivity vs. (1 − specificity) across all possible cut-offs; AUC quantifies overall test performance
Post-test ProbabilityPPV and NPV (prevalence-dependent)Fagan nomogram uses pre-test probability × LR to calculate post-test probability (Bayesian reasoning)
Risk QuantificationRR, OR, AR for single exposuresMultivariable regression (logistic, Cox) adjusts for confounders and yields adjusted OR or hazard ratios
Population ImpactAttributable risk in the exposedPopulation attributable risk (PAR) estimates disease burden in the total population attributable to a given exposure

The Receiver Operating Characteristic (ROC) curve is particularly important. It plots sensitivity (y-axis) against 1 − specificity (x-axis) for every possible test threshold. The area under the ROC curve (AUC) ranges from 0.5 (no discriminatory power, equivalent to flipping a coin) to 1.0 (perfect discrimination). When comparing two screening tests for the same disease, the test with the higher AUC is generally superior. You may encounter Step 1 questions asking you to identify which point on an ROC curve optimizes the trade-off between sensitivity and specificity for a given clinical scenario.

Practice Problems

PROBLEM 1CONCEPTUAL
A new screening test for colorectal cancer has 95% sensitivity and 80% specificity. If this test is applied to a low-prevalence population (0.5%) versus a high-prevalence population (10%), how will the positive predictive value (PPV) differ between the two populations, and why? Explain conceptually without performing exact calculations.
PROBLEM 2BASIC CALCULATION
A cohort study of 1,000 smokers and 1,000 non-smokers tracked lung cancer incidence over 20 years. Among smokers, 180 developed lung cancer. Among non-smokers, 20 developed lung cancer. Calculate the relative risk (RR) and the attributable risk (AR).
PROBLEM 3INTERMEDIATE
A case-control study examines the association between oral contraceptive (OC) use and deep vein thrombosis (DVT). Among 200 women with DVT (cases), 80 used OCs. Among 400 women without DVT (controls), 100 used OCs. Calculate the odds ratio. Can you compute a relative risk from this study design? Explain your answer.
PROBLEM 4APPLIED
A hospital implements a rapid antigen test for influenza with 70% sensitivity and 98% specificity. During peak flu season, the prevalence of influenza among symptomatic patients presenting to the ED is 25%. Calculate the PPV and NPV using a hypothetical population of 1,000 patients. Should a negative result be trusted to rule out influenza in this setting?
PROBLEM 5CRITICAL THINKING
A public health official argues that a new cancer screening program reduced 5-year mortality by 20% compared to unscreened populations. A biostatistician counters that this improvement could be entirely explained by lead-time bias and length-time bias without any true mortality benefit. Explain both biases and describe what study design feature would best address this criticism.

Summary

Risk measures and screening metrics all derive from the 2 × 2 contingency table. Relative risk (RR) is the ratio of incidence in exposed versus unexposed groups and is used in cohort studies and RCTs. The odds ratio (OR) is the primary measure of association in case-control studies and approximates RR when disease prevalence is low. Attributable risk (AR) is the absolute risk difference, and its reciprocal yields the number needed to treat (NNT) or number needed to harm (NNH).

For screening tests, sensitivity and specificity are intrinsic test properties: remember SnNOut (sensitive test, negative result rules out) and SpPIn (specific test, positive result rules in). PPV and NPV depend on disease prevalence: higher prevalence raises PPV and lowers NPV. Advanced extensions include likelihood ratios and ROC curves for refined diagnostic reasoning. Be alert to lead-time bias and length-time bias when interpreting screening program outcomes.

Varsity Tutors • USMLE Step 1 • Risk Measures And Screening