GENETICS • MOLECULAR GENETICS TECHNIQUES & GENOMICS

SNPs & Genetic Variation — SNPs and genetic variation concepts

Tiny single-letter DNA changes that make every person unique and shape human health.

Historical Context & Motivation

Have you ever wondered why you look different from your classmates, even though you are all human? The answer lies deep inside your DNA — the long molecule that stores the instructions for building and running your body. For centuries, scientists knew that traits were inherited, but they didn't understand exactly how tiny differences in DNA could lead to so much variety among people.

Once the Human Genome Project finished mapping out all 3 billion letters of human DNA, researchers realized something amazing: about 99.9% of DNA is identical from person to person. The remaining 0.1% contains the differences that make each of us unique. The most common type of difference is a single nucleotide polymorphism, or SNP (pronounced "snip"). Understanding SNPs has become one of the most important goals in modern genetics.

1953
DNA Structure Discovered
James Watson and Francis Crick, building on Rosalind Franklin's X-ray data, reveal the double-helix shape of DNA — setting the stage for reading its letters.
1990
Human Genome Project Begins
An international team starts the ambitious effort to read every one of the 3 billion DNA base pairs in a human cell.
2001
First Draft of the Human Genome
The first near-complete sequence is published, revealing that humans share 99.9% of their DNA. Scientists begin cataloging the tiny differences — SNPs.
2002
HapMap Project Launched
The International HapMap Project maps common SNP patterns across different populations, creating a powerful tool for studying how genetic variation connects to health.
2010s–Now
GWAS & Personal Genomics
Genome-wide association studies (GWAS) scan millions of SNPs to link them to diseases. Consumer DNA tests let anyone explore their own SNP data.

The central question that drove all of this research is: How do such tiny changes in DNA lead to big differences in appearance, health, and even disease risk? That is the question this lesson will help you answer.

Core Principles & Definitions

Before diving deeper, let's lock down the key ideas. DNA is made of four chemical building blocks called nucleotides, represented by the letters A (adenine), T (thymine), C (cytosine), and G (guanine). A SNP occurs when a single letter in the DNA sequence differs between individuals. For example, at one specific spot in the genome, most people might have the letter A, while some people have a G instead.

1

What Is a SNP?

A single nucleotide polymorphism is a one-letter change at a specific position in DNA. To count as a SNP, the less common version (called the minor allele) must appear in at least 1% of the population.
2

Alleles

Each version of the nucleotide at a SNP position is called an allele. Most SNPs are biallelic, meaning they have only two possible letters — for example, A or G.
3

Genotype

Because you inherit one copy of each chromosome from each parent, you carry two alleles at every SNP. Your combination — AA, AG, or GG — is your genotype at that position.
4

Minor Allele Frequency (MAF)

The minor allele frequency is the proportion of all alleles in a population that are the less common type. A MAF of 0.30 means 30% of all copies carry the minor allele.
5

Phenotype Connection

A phenotype is a trait you can observe — like eye color or height. Some SNPs influence phenotypes directly; others have no visible effect at all.
KEY TAKEAWAY
Think of your DNA as a very long book — about 3 billion letters long. A SNP is like a single typo on one page. Most of the time the typo doesn't change the story at all. But once in a while, that one changed letter changes a key word and alters the meaning of the sentence. That's how a single SNP can sometimes change how your body works.

Visual Explanation — Seeing a SNP

The diagram below shows a short stretch of DNA from two different people. Notice that the sequences are nearly identical — the only difference is at one position, highlighted in color. Person A carries a C (cytosine) at that spot, while Person B carries a T (thymine). This single-letter swap is exactly what a SNP looks like.

Two DNA sequences are shown letter by letter. The highlighted position is where Person A has a C and Person B has a T. Every other letter is the same. This single-nucleotide change is a SNP.

In the diagram above, the boxed letters are the SNP. Everything else is identical between the two people. Scientists estimate that there are roughly 4 to 5 million SNPs in any single person's genome when compared to a reference sequence. Across all of humanity, over 600 million SNP positions have been cataloged.

How SNPs Arise & Allele Frequency

SNPs are created by point mutations — errors that happen when DNA is copied during cell division. Each time a cell divides, it must duplicate all 3 billion letters. The copying machinery is very accurate, but it still makes roughly one mistake per every 100 million letters. Most of these mistakes are repaired immediately, but a few slip through. If a mutation occurs in a reproductive cell (sperm or egg), it can be passed to the next generation and eventually spread through a population.

Scientists use a simple formula to describe how common each version of a SNP is. For a SNP with two alleles, we label the more common one p and the less common one q. Because every allele in the population must be one or the other, their frequencies always add up to 1.

ALLELE FREQUENCY RULE
p + q = 1
p = frequency of the major (more common) allele; q = frequency of the minor (less common) allele. Frequencies are expressed as decimals between 0 and 1.

If we know one frequency, we can find the other. For example, if the major allele has a frequency of 0.75, then the minor allele frequency is 1 − 0.75 = 0.25. This means that 25% of all copies of that chromosome in the population carry the minor allele.

EXPECTED GENOTYPE FREQUENCIES (HARDY-WEINBERG)
p² + 2pq + q² = 1
p² = expected frequency of homozygous major genotype (e.g., AA); 2pq = expected frequency of heterozygous genotype (e.g., AG); q² = expected frequency of homozygous minor genotype (e.g., GG). This equation assumes random mating and is called the Hardy-Weinberg equation.
💡 Why 1% Matters
A DNA change is only called a SNP if the minor allele frequency is at least 1% in the population. Changes that are rarer than 1% are typically called rare variants instead. This threshold helps scientists focus on variations that are common enough to study in large groups.

Types of SNPs & Their Effects

Not all SNPs are created equal. Their impact depends on where in the genome they occur. Some SNPs sit inside genes (the sections of DNA that code for proteins), while others sit in non-coding regions between genes. The diagram below classifies the major types and their potential effects.

A classification tree showing the major types of SNPs. Coding-region SNPs can be synonymous (silent) or non-synonymous (changing the protein). Non-coding SNPs may affect gene regulation or have no known effect.
Summary of SNP types, their genomic locations, and effects
SNP TypeLocationEffect on ProteinReal-World Example
SynonymousInside a gene (exon)None — same amino acid producedMany codons have this kind of "wobble" at the third position
MissenseInside a gene (exon)One amino acid swapped for anotherSickle-cell disease: glutamic acid → valine in hemoglobin
NonsenseInside a gene (exon)Creates a premature stop codon — shorter proteinSome forms of cystic fibrosis involve early stop codons
RegulatoryNear a gene (promoter / enhancer)Protein is normal but made in wrong amountLactase persistence: ability to digest milk as an adult
IntergenicBetween genesUsually noneUsed as genetic markers in ancestry testing

Worked Example — Calculating Allele Frequencies

Imagine a class of 50 students who have been genotyped at a specific SNP. The SNP has two alleles: A (major) and G (minor). The genotype counts are: 30 students are AA, 16 students are AG, and 4 students are GG. Let's find the allele frequencies and check them with the Hardy-Weinberg equation.

Finding Allele Frequencies from Genotype Data
1
Step 1 — Count Total AllelesEach person carries 2 alleles (one from each parent). With 50 students, the total number of alleles is 50 × 2 = 100.
Total alleles = 100
2
Step 2 — Count Each AlleleAA students contribute 2 A alleles each: 30 × 2 = 60 A alleles. AG students contribute 1 A and 1 G each: 16 × 1 = 16 A alleles and 16 × 1 = 16 G alleles. GG students contribute 2 G alleles each: 4 × 2 = 8 G alleles. Total A alleles = 60 + 16 = 76. Total G alleles = 16 + 8 = 24.
A count = 76; G count = 24
3
Step 3 — Calculate FrequenciesFrequency of A (p) = 76 ÷ 100 = 0.76. Frequency of G (q) = 24 ÷ 100 = 0.24. Check: p + q = 0.76 + 0.24 = 1.00 ✓
p = 0.76, q = 0.24
4
Step 4 — Predict Genotype Frequencies (Hardy-Weinberg)Expected AA frequency = p² = 0.76² = 0.5776. Expected AG frequency = 2pq = 2 × 0.76 × 0.24 = 0.3648. Expected GG frequency = q² = 0.24² = 0.0576. Check: 0.5776 + 0.3648 + 0.0576 = 1.0000 ✓
Expected: AA = 0.578, AG = 0.365, GG = 0.058
5
Step 5 — Compare Observed vs. ExpectedObserved frequencies: AA = 30/50 = 0.60, AG = 16/50 = 0.32, GG = 4/50 = 0.08. The observed values are close to the expected Hardy-Weinberg values, suggesting the population is roughly in equilibrium for this SNP. Small differences could be due to random sampling in a class of only 50.
Observed ≈ Expected → population is near Hardy-Weinberg equilibrium

Applications, Strengths & Limitations of SNP Analysis

SNP data has transformed many areas of science and medicine. Doctors use SNPs to predict drug responses (a field called pharmacogenomics), forensic scientists use them to identify individuals, and anthropologists use them to trace human migration patterns. However, SNP analysis has important limitations too.

Strengths and limitations of SNP-based genetic analysis
StrengthsLimitations
Abundant: millions of SNPs across the genome give dense coverageMost SNPs have no effect — finding the important ones requires huge studies
Cheap and fast to genotype using SNP microarray chipsSNP chips only test known SNPs — they miss brand-new or very rare variants
Stable: SNPs mutate slowly, so they are reliable markers over generationsAssociation ≠ causation: a linked SNP may be near a causal variant, not the cause itself
Useful for ancestry analysis, disease risk, and drug responseComplex traits (height, intelligence) involve thousands of SNPs, each with tiny effects
Can be combined with other data (gene expression, environment) for deeper insightsMost GWAS databases over-represent people of European descent, limiting applicability
KEY TAKEAWAY
Think of SNP analysis like a weather forecast. It uses millions of data points to make a prediction, and those predictions are useful — but they are never 100% certain. A SNP can raise or lower your probability of a trait, just like weather data raises or lowers the probability of rain. Your genes are just one ingredient; environment, lifestyle, and chance also play big roles.

SNPs vs. Other Types of Genetic Variation

SNPs are the most common form of genetic variation, but they are not the only kind. Other types of variation involve more than a single nucleotide. Understanding how SNPs compare to these other forms helps you see the full picture of human genetic diversity.

Comparison of SNPs with other forms of genetic variation
FeatureSNPsInsertions / Deletions (Indels)Copy Number Variants (CNVs)
Size of change1 nucleotide1 to ~1,000 nucleotides added or removedThousands to millions of nucleotides duplicated or deleted
Frequency in genome~4–5 million per person~500,000–800,000 per person~1,000+ per person
Detection methodSNP microarray; sequencingSequencingArray comparative genomic hybridization; sequencing
Typical effectUsually neutral; some affect traits or disease riskCan shift the reading frame of a gene, often harmfulCan change gene dosage — more or fewer copies of a gene
Examplers1426654 — affects skin pigmentationΔF508 deletion in CFTR gene → cystic fibrosisExtra copies of AMY1 gene → better starch digestion

As sequencing technology gets cheaper and faster, scientists are moving beyond SNP chips to whole-genome sequencing, which reads every single letter of a person's DNA. This approach captures SNPs, indels, CNVs, and structural variants all at once. In the future, your doctor may use your full genome sequence — not just a list of SNPs — to guide your healthcare.

Practice Problems

PROBLEM 1CONCEPTUAL
In your own words, explain the difference between a SNP and a mutation. Why do scientists use the term "polymorphism" instead of "mutation" when talking about SNPs?
PROBLEM 2BASIC CALCULATION
A population of 200 people is genotyped at a SNP with alleles C and T. The results are: 120 CC, 64 CT, and 16 TT. Calculate the frequency of the C allele (p) and the T allele (q).
PROBLEM 3INTERMEDIATE
Using the allele frequencies from Problem 2 (p = 0.76, q = 0.24), predict the expected genotype frequencies under Hardy-Weinberg equilibrium. Then calculate the expected number of each genotype in the 200-person population and compare to the observed numbers (120 CC, 64 CT, 16 TT).
PROBLEM 4APPLIED
A pharmaceutical company discovers that patients with the GG genotype at a particular SNP respond poorly to a new drug, while AA and AG patients respond well. In a population where the G allele frequency is 0.15, what percentage of people would be expected to have the GG genotype? If 10,000 patients need this drug, approximately how many would be poor responders?
PROBLEM 5CRITICAL THINKING
A consumer DNA testing company reports that you carry a SNP associated with a 1.3× increased risk for Type 2 diabetes. The average lifetime risk for Type 2 diabetes is about 10%. Does this SNP mean you will definitely develop diabetes? Discuss at least three reasons why a single SNP result should be interpreted with caution.

Lesson Summary

A single nucleotide polymorphism (SNP) is a one-letter change in the DNA sequence that occurs at a specific position and is found in at least 1% of the population. Humans share about 99.9% of their DNA, and SNPs account for the most common type of variation in the remaining 0.1%. Each person carries roughly 4 to 5 million SNPs. SNPs can be classified as synonymous, missense, nonsense, regulatory, or intergenic depending on their location and effect on proteins.

The frequency of each allele at a SNP position can be calculated using simple counting (p + q = 1), and the Hardy-Weinberg equation (p² + 2pq + q² = 1) predicts expected genotype frequencies under random mating. SNP analysis powers modern tools like GWAS, pharmacogenomics, and ancestry testing, but results must be interpreted carefully — most complex traits depend on thousands of SNPs plus environmental factors. As technology advances toward whole-genome sequencing, scientists will capture an even fuller picture of human genetic variation.

Varsity Tutors • Genetics • SNPs & Genetic Variation