GENETICS • MOLECULAR GENETICS TECHNIQUES & GENOMICS

DNA Sequencing

Unlocking the genetic code letter by letter to understand life's blueprint.

Historical Context & Motivation

Imagine you have a recipe book with billions of letters, and you need to read every single one in the correct order. That is essentially what DNA sequencing does — it figures out the exact order of the chemical "letters" (called nucleotide bases) that make up a strand of DNA. Knowing this order helps scientists understand how organisms grow, why diseases happen, and even how different species are related to one another.

Before DNA sequencing was invented, scientists knew that DNA carried genetic information, but they had no way to read it precisely. The race to develop reliable sequencing methods transformed biology from a descriptive science into one that could decode life at its most fundamental level.

1953
DNA Double Helix Discovered
James Watson and Francis Crick, building on Rosalind Franklin's X-ray crystallography data, revealed the double-helix structure of DNA. This showed scientists what they needed to "read" — a twisted ladder of paired bases.
1977
Sanger Sequencing Invented
Frederick Sanger developed a method using special chain-terminating nucleotides to read DNA one base at a time. This Sanger sequencing technique became the gold standard for decades.
1990
Human Genome Project Launched
An international effort began to sequence the entire human genome — roughly 3.2 billion base pairs. It was the largest collaborative biology project in history.
2003
Human Genome Completed
After 13 years and approximately $2.7 billion, scientists published a nearly complete map of human DNA. This achievement opened the door to personalized medicine and modern genomics.
2005–Present
Next-Generation Sequencing (NGS)
New technologies emerged that could read millions of DNA fragments at once, slashing the cost from billions of dollars to under $1,000 per genome. Next-generation sequencing made DNA analysis fast, affordable, and widely available.

The central question that drove all of this work was simple but powerful: How can we read the exact sequence of bases in DNA quickly and accurately? Every method we will explore in this lesson is an answer to that question, and each generation of technology has made the answer faster, cheaper, and more precise.

Core Principles of DNA Sequencing

Before diving into specific methods, you need to understand a few building blocks. DNA is made of four nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). These bases pair up — A always pairs with T, and C always pairs with G. This is called complementary base pairing. Every sequencing method takes advantage of these pairing rules in some way.

1

Template & Primer

Sequencing starts with a single-stranded DNA template (the strand you want to read). A short piece of DNA called a primer binds to the template to give the copying enzyme a starting point.
2

DNA Polymerase

An enzyme called DNA polymerase reads the template strand and adds matching bases one at a time. It builds a new complementary strand, following the base-pairing rules.
3

Chain Termination

In Sanger sequencing, special modified bases called dideoxynucleotides (ddNTPs) are mixed in. When one is added, the growing chain stops. This creates fragments of different lengths, each ending with a known base.
4

Detection & Reading

Each ddNTP carries a different fluorescent dye (one color per base). A laser excites the dyes, and a detector reads the color of each fragment to determine the base at that position.
5

Assembly & Alignment

Since most genomes are too long to read in one pass, DNA is broken into smaller pieces, sequenced separately, and then assembled by computer software that lines up overlapping regions to reconstruct the full sequence.
KEY TAKEAWAY
Think of DNA sequencing like solving a giant jigsaw puzzle. First, you photocopy the puzzle picture onto many sheets and cut each sheet into different-sized strips (fragmentation). Then you read the last letter on each strip (detection). Finally, you overlap the strips to reconstruct the whole message (assembly). Every sequencing method follows this basic pattern: copy, fragment, read, and assemble.

Visualizing the Sanger Sequencing Process

The diagram below walks you through the core steps of Sanger sequencing (also called the chain-termination method). Follow the flow from top to bottom to see how a single strand of DNA gets copied, fragmented, separated, and finally read by a detector.

This diagram shows the four main steps of Sanger sequencing. Step 1 sets up the template and primer. Step 2 uses DNA polymerase with a mix of normal and chain-terminating bases. Step 3 separates fragments by size. Step 4 reads the fluorescent signal to produce a chromatogram — the colored peaks that reveal the DNA sequence.

Notice that the key to Sanger sequencing is the random incorporation of ddNTPs. Because the chain-terminating bases are mixed at a low concentration with normal bases, the polymerase sometimes grabs a ddNTP and sometimes grabs a normal one. Over millions of reactions, this creates a complete set of fragments ending at every position in the sequence. When those fragments are lined up shortest to longest, you can read the bases in order — just like reading a sentence one letter at a time.

How Sequencing Works — The Molecular Details

Let's dig deeper into what happens at the molecular level during DNA sequencing. Understanding these details will help you see why the method works and what its limits are.

Normal vs. Chain-Terminating Nucleotides

A normal nucleotide (called a deoxyribonucleotide triphosphate, or dNTP) has a hydroxyl group (−OH) on its 3ʼ carbon. This −OH group is essential because it forms the bond with the next nucleotide in the chain. A dideoxyribonucleotide triphosphate (ddNTP) is missing that −OH group — it has just a hydrogen (−H) instead. Without the −OH, no new bond can form, and the chain stops growing.

CHAIN TERMINATION PRINCIPLE
dNTP (3ʼ −OH) → chain continues | ddNTP (3ʼ −H) → chain STOPS
dNTP = normal base with a 3ʼ hydroxyl group; ddNTP = modified base without the 3ʼ hydroxyl group. The missing −OH prevents the phosphodiester bond from forming.

Sequencing Coverage & Read Length

In genomics, scientists talk about coverage — how many times, on average, each base in a genome has been sequenced. Higher coverage means more confidence that the sequence is correct. The basic formula for average coverage is:

SEQUENCING COVERAGE
Coverage (C) = (N × L) ÷ G
C = average coverage (how many times each base is read), N = total number of reads (fragments sequenced), L = average read length (in base pairs), G = genome size (in base pairs). Most projects aim for 30× coverage or higher.

PCR Amplification Before Sequencing

Most sequencing methods require millions of copies of the DNA to get a strong enough signal. The polymerase chain reaction (PCR) is used to amplify DNA before sequencing. Each PCR cycle doubles the number of copies, so after n cycles you have approximately 2n copies of the target region.

PCR AMPLIFICATION
Number of copies ≈ 2ⁿ
n = number of PCR cycles. For example, after 30 cycles: 2³⁰ ≈ 1.07 × 10⁹ copies (over one billion copies from a single DNA molecule).
🔬 Why Coverage Matters
DNA sequencing is not perfect — errors can occur during any read. If a base is sequenced only once, you can't tell whether the result is real or a mistake. With 30× coverage, each base is read about 30 separate times, so scientists can compare reads and confidently identify the correct base. Think of it like asking 30 people to spell a word — if 29 say "genome" and one says "gnome," you know the right answer.

Types of DNA Sequencing Methods

Since Sanger's original method, many new sequencing technologies have been developed. They fall into three broad generations, each offering different advantages in speed, cost, and read length. The diagram below compares the major methods.

This comparison chart shows three generations of sequencing technology. First-generation (Sanger) offers the highest per-read accuracy but is slow and expensive. Second-generation (NGS) massively increased throughput and slashed costs. Third-generation technologies read extremely long fragments in real time, making them ideal for complex genome regions.

How Illumina Sequencing Works

The most widely used sequencing platform today is Illumina sequencing, which uses a method called sequencing by synthesis (SBS). DNA is fragmented, attached to a glass surface (called a flow cell), and amplified into tiny clusters. Then, fluorescently labeled bases are added one at a time. After each base is added, a camera takes a picture to record which base was incorporated at each cluster. This process repeats for every cycle, building up the sequence one letter at a time across millions of clusters simultaneously.

How Nanopore Sequencing Works

Oxford Nanopore sequencing takes a completely different approach. A single strand of DNA is threaded through a tiny protein pore (a nanopore) embedded in a membrane. As each base passes through the pore, it changes the electrical current flowing through the pore by a slightly different amount. A sensor detects these current changes in real time and translates them into the DNA sequence. This method can read extremely long fragments — sometimes over 100,000 bases in a single read — and the device can be as small as a USB drive.

Worked Example: Calculating Sequencing Coverage

Let's work through a realistic problem to see how scientists plan a sequencing experiment. We'll use the coverage formula from Section 4.

Calculating Coverage for a Bacterial Genome
1
Step 1 — Identify the Given ValuesA researcher wants to sequence a bacterial genome that is 5,000,000 base pairs (5 × 10⁶ bp) long. She is using an Illumina sequencer that produces 10,000,000 reads (10⁷ reads), each with an average read length of 150 bp. She wants to know: what is the average coverage?
G = 5 × 10⁶ bp, N = 10⁷ reads, L = 150 bp
2
Step 2 — Write the Coverage FormulaWe use the coverage formula: C = (N × L) ÷ G. This tells us how many times, on average, each base in the genome will be sequenced.
C = (N × L) ÷ G
3
Step 3 — Substitute the ValuesPlug in the numbers: C = (10,000,000 × 150) ÷ 5,000,000. First, multiply the numerator: 10,000,000 × 150 = 1,500,000,000 (which is 1.5 × 10⁹).
Numerator = 1.5 × 10⁹
4
Step 4 — Divide to Find CoverageNow divide: C = 1,500,000,000 ÷ 5,000,000 = 300. This means each base in the genome is read, on average, 300 times.
C = 300× coverage
5
Step 5 — Interpret the ResultA coverage of 300× is very high — much more than the typical 30× target. This means the researcher has more data than she needs for a high-confidence sequence. She could either use a smaller run to save cost, or use the extra data for very precise variant detection. High coverage is especially useful when searching for rare mutations that might only appear in a small fraction of cells.
300× is well above the 30× standard — excellent data quality

Comparing Sequencing Methods: Strengths & Limitations

No single sequencing method is perfect for every situation. The best choice depends on what you need: high accuracy, long reads, low cost, or portability. The table below summarizes the key trade-offs between the three major approaches.

Comparison of three generations of DNA sequencing technology
FeatureSanger (1st Gen)Illumina (2nd Gen)Nanopore (3rd Gen)
Read Length700–1,000 bp (long)75–300 bp (short)10,000–100,000+ bp (very long)
Accuracy99.99% (highest)99.9% (very high)95–99% (improving)
ThroughputLow (one fragment at a time)Very high (millions of reads)Moderate
Cost per GenomeVery expensive$200–$1,000 (cheapest)$500–$1,500
SpeedHours to days1–3 daysReal-time (minutes to hours)
PortabilityLab-basedLab-basedHandheld (MinION)
Main LimitationSlow, costly for whole genomesShort reads; struggles with repetitive DNAHigher error rate per read
KEY TAKEAWAY
Choosing a sequencing method is like choosing a vehicle for a trip. Sanger sequencing is like a luxury car — precise and reliable, but slow and expensive for long journeys. Illumina is like a fleet of buses — it can move millions of passengers (reads) at once for a low fare, but each trip is short. Nanopore is like a helicopter — it can cover long distances in one go and go anywhere, but the ride is a bit bumpier. Many modern projects combine methods to get the best of all worlds.

Applications & Connection to Advanced Genomics

DNA sequencing is not just a laboratory technique — it has transformed fields from medicine to agriculture to criminal justice. Understanding the basics you've learned here prepares you for more advanced topics in genomics, bioinformatics, and genetic engineering.

Real-world applications of DNA sequencing and their connections to advanced genomics
Application AreaHow DNA Sequencing Is UsedAdvanced Connection
Medical DiagnosisIdentify genetic mutations that cause diseases like cystic fibrosis or sickle cell anemiaLeads to whole-exome sequencing, pharmacogenomics, and precision medicine
Cancer ResearchCompare tumor DNA to normal DNA to find driver mutationsConnects to cancer genomics, liquid biopsy, and targeted therapy
Forensic ScienceMatch crime scene DNA to suspects using short tandem repeat (STR) analysisAdvances to forensic genomics and genetic genealogy databases
Evolutionary BiologyCompare DNA sequences across species to build family trees of lifeUnderpins phylogenomics, ancient DNA studies, and molecular clocks
AgricultureIdentify genes for drought resistance, disease resistance, or higher yield in cropsConnects to genomic-assisted breeding and CRISPR gene editing
Pandemic ResponseSequence pathogen genomes (e.g., SARS-CoV-2) to track variants and develop vaccinesConnects to metagenomics and real-time epidemiological surveillance

As sequencing costs continue to fall and new technologies emerge, the applications of DNA sequencing will only grow. Scientists are already working on sequencing entire genomes in minutes, reading DNA modifications like epigenetic marks (chemical tags that control gene activity without changing the DNA letters), and even storing digital data in synthetic DNA. The foundation you've built in this lesson — understanding templates, polymerases, base pairing, and coverage — will serve you in any of these advanced areas.

🚀 Looking Ahead
If DNA sequencing tells us what letters are in the genetic code, then bioinformatics is the field that helps us understand what those letters mean. After sequencing, computers compare sequences, predict protein structures, and identify patterns. As you advance in genetics, bioinformatics will become your most important partner.

Practice Problems

PROBLEM 1CONCEPTUAL
In Sanger sequencing, why are dideoxynucleotides (ddNTPs) essential to the process? Explain what would happen if only normal nucleotides (dNTPs) were used.
PROBLEM 2BASIC CALCULATION
A lab runs 20 PCR cycles to amplify a DNA sample before sequencing. Starting from a single copy, approximately how many copies of the target DNA will be produced?
PROBLEM 3INTERMEDIATE
A sequencing experiment on a 3,200,000,000 bp (3.2 × 10⁹ bp) human genome produces 500,000,000 reads, each 300 bp long. Calculate the average coverage. Is this sufficient for a clinical diagnostic sequencing project that requires at least 30× coverage?
PROBLEM 4APPLIED
During a disease outbreak in a remote village, health workers need to quickly identify whether the pathogen is a new variant of a known virus. They have limited electricity and no access to a full laboratory. Which generation of sequencing technology would be most appropriate, and why? What trade-offs would they accept?
PROBLEM 5CRITICAL THINKING
A genome has a large region of highly repetitive DNA — the same 500 bp sequence is repeated 200 times in a row (100,000 bp total). Explain why short-read sequencing (like Illumina with 150 bp reads) would struggle to correctly assemble this region. How could long-read sequencing solve this problem? What strategy might a research team use if they only have access to short-read technology?

DNA Sequencing — Key Concepts Review

DNA sequencing is the process of determining the exact order of nucleotide bases (A, T, C, G) in a strand of DNA. The foundational method, Sanger sequencing, uses chain-terminating dideoxynucleotides (ddNTPs) to create fragments of different lengths, which are separated and read by fluorescence detection. All sequencing methods rely on DNA polymerase, complementary base pairing, and PCR amplification (producing 2ⁿ copies in n cycles) to generate enough material for accurate reading.

Modern sequencing spans three generations: first-generation (Sanger) for high-accuracy single reads, second-generation (Illumina/NGS) for massively parallel short reads at low cost, and third-generation (Nanopore/PacBio) for ultra-long reads in real time. Scientists calculate sequencing coverage using the formula C = (N × L) ÷ G, aiming for at least 30× for reliable results. Applications range from medical diagnosis and forensics to pandemic surveillance and evolutionary biology, making DNA sequencing one of the most transformative tools in modern science.

Varsity Tutors • Genetics • DNA Sequencing