Historical Context & Motivation
Have you ever wondered why you look different from your classmates, even though you are all human? The answer lies deep inside your DNA — the long molecule that stores the instructions for building and running your body. For centuries, scientists knew that traits were inherited, but they didn't understand exactly how tiny differences in DNA could lead to so much variety among people.
Once the Human Genome Project finished mapping out all 3 billion letters of human DNA, researchers realized something amazing: about 99.9% of DNA is identical from person to person. The remaining 0.1% contains the differences that make each of us unique. The most common type of difference is a single nucleotide polymorphism, or SNP (pronounced "snip"). Understanding SNPs has become one of the most important goals in modern genetics.
The central question that drove all of this research is: How do such tiny changes in DNA lead to big differences in appearance, health, and even disease risk? That is the question this lesson will help you answer.
Core Principles & Definitions
Before diving deeper, let's lock down the key ideas. DNA is made of four chemical building blocks called nucleotides, represented by the letters A (adenine), T (thymine), C (cytosine), and G (guanine). A SNP occurs when a single letter in the DNA sequence differs between individuals. For example, at one specific spot in the genome, most people might have the letter A, while some people have a G instead.
What Is a SNP?
Alleles
Genotype
Minor Allele Frequency (MAF)
Phenotype Connection
Visual Explanation — Seeing a SNP
The diagram below shows a short stretch of DNA from two different people. Notice that the sequences are nearly identical — the only difference is at one position, highlighted in color. Person A carries a C (cytosine) at that spot, while Person B carries a T (thymine). This single-letter swap is exactly what a SNP looks like.
In the diagram above, the boxed letters are the SNP. Everything else is identical between the two people. Scientists estimate that there are roughly 4 to 5 million SNPs in any single person's genome when compared to a reference sequence. Across all of humanity, over 600 million SNP positions have been cataloged.
How SNPs Arise & Allele Frequency
SNPs are created by point mutations — errors that happen when DNA is copied during cell division. Each time a cell divides, it must duplicate all 3 billion letters. The copying machinery is very accurate, but it still makes roughly one mistake per every 100 million letters. Most of these mistakes are repaired immediately, but a few slip through. If a mutation occurs in a reproductive cell (sperm or egg), it can be passed to the next generation and eventually spread through a population.
Scientists use a simple formula to describe how common each version of a SNP is. For a SNP with two alleles, we label the more common one p and the less common one q. Because every allele in the population must be one or the other, their frequencies always add up to 1.
If we know one frequency, we can find the other. For example, if the major allele has a frequency of 0.75, then the minor allele frequency is 1 − 0.75 = 0.25. This means that 25% of all copies of that chromosome in the population carry the minor allele.
Types of SNPs & Their Effects
Not all SNPs are created equal. Their impact depends on where in the genome they occur. Some SNPs sit inside genes (the sections of DNA that code for proteins), while others sit in non-coding regions between genes. The diagram below classifies the major types and their potential effects.
| SNP Type | Location | Effect on Protein | Real-World Example |
|---|---|---|---|
| Synonymous | Inside a gene (exon) | None — same amino acid produced | Many codons have this kind of "wobble" at the third position |
| Missense | Inside a gene (exon) | One amino acid swapped for another | Sickle-cell disease: glutamic acid → valine in hemoglobin |
| Nonsense | Inside a gene (exon) | Creates a premature stop codon — shorter protein | Some forms of cystic fibrosis involve early stop codons |
| Regulatory | Near a gene (promoter / enhancer) | Protein is normal but made in wrong amount | Lactase persistence: ability to digest milk as an adult |
| Intergenic | Between genes | Usually none | Used as genetic markers in ancestry testing |
Worked Example — Calculating Allele Frequencies
Imagine a class of 50 students who have been genotyped at a specific SNP. The SNP has two alleles: A (major) and G (minor). The genotype counts are: 30 students are AA, 16 students are AG, and 4 students are GG. Let's find the allele frequencies and check them with the Hardy-Weinberg equation.
Applications, Strengths & Limitations of SNP Analysis
SNP data has transformed many areas of science and medicine. Doctors use SNPs to predict drug responses (a field called pharmacogenomics), forensic scientists use them to identify individuals, and anthropologists use them to trace human migration patterns. However, SNP analysis has important limitations too.
| Strengths | Limitations |
|---|---|
| Abundant: millions of SNPs across the genome give dense coverage | Most SNPs have no effect — finding the important ones requires huge studies |
| Cheap and fast to genotype using SNP microarray chips | SNP chips only test known SNPs — they miss brand-new or very rare variants |
| Stable: SNPs mutate slowly, so they are reliable markers over generations | Association ≠ causation: a linked SNP may be near a causal variant, not the cause itself |
| Useful for ancestry analysis, disease risk, and drug response | Complex traits (height, intelligence) involve thousands of SNPs, each with tiny effects |
| Can be combined with other data (gene expression, environment) for deeper insights | Most GWAS databases over-represent people of European descent, limiting applicability |
SNPs vs. Other Types of Genetic Variation
SNPs are the most common form of genetic variation, but they are not the only kind. Other types of variation involve more than a single nucleotide. Understanding how SNPs compare to these other forms helps you see the full picture of human genetic diversity.
| Feature | SNPs | Insertions / Deletions (Indels) | Copy Number Variants (CNVs) |
|---|---|---|---|
| Size of change | 1 nucleotide | 1 to ~1,000 nucleotides added or removed | Thousands to millions of nucleotides duplicated or deleted |
| Frequency in genome | ~4–5 million per person | ~500,000–800,000 per person | ~1,000+ per person |
| Detection method | SNP microarray; sequencing | Sequencing | Array comparative genomic hybridization; sequencing |
| Typical effect | Usually neutral; some affect traits or disease risk | Can shift the reading frame of a gene, often harmful | Can change gene dosage — more or fewer copies of a gene |
| Example | rs1426654 — affects skin pigmentation | ΔF508 deletion in CFTR gene → cystic fibrosis | Extra copies of AMY1 gene → better starch digestion |
As sequencing technology gets cheaper and faster, scientists are moving beyond SNP chips to whole-genome sequencing, which reads every single letter of a person's DNA. This approach captures SNPs, indels, CNVs, and structural variants all at once. In the future, your doctor may use your full genome sequence — not just a list of SNPs — to guide your healthcare.
Practice Problems
Lesson Summary
A single nucleotide polymorphism (SNP) is a one-letter change in the DNA sequence that occurs at a specific position and is found in at least 1% of the population. Humans share about 99.9% of their DNA, and SNPs account for the most common type of variation in the remaining 0.1%. Each person carries roughly 4 to 5 million SNPs. SNPs can be classified as synonymous, missense, nonsense, regulatory, or intergenic depending on their location and effect on proteins.
The frequency of each allele at a SNP position can be calculated using simple counting (p + q = 1), and the Hardy-Weinberg equation (p² + 2pq + q² = 1) predicts expected genotype frequencies under random mating. SNP analysis powers modern tools like GWAS, pharmacogenomics, and ancestry testing, but results must be interpreted carefully — most complex traits depend on thousands of SNPs plus environmental factors. As technology advances toward whole-genome sequencing, scientists will capture an even fuller picture of human genetic variation.