Historical Context & Motivation
For most of human history, the mechanism of heredity remained a mystery. Farmers and breeders understood that offspring resemble their parents, but no one could explain how traits pass from one generation to the next at the molecular level. The search for the molecule responsible for inheritance took more than a century of experiments, controversies, and breakthroughs. Each discovery built on the last, gradually revealing that a single molecule — deoxyribonucleic acid (DNA) — carries the instructions for building and operating every living organism.
These discoveries raised a fundamental question: how does the specific sequence of chemical units in DNA encode the information needed to build thousands of different proteins? This lesson answers that question by tracing the flow of genetic information from the base sequence of DNA to the amino acid sequence of a protein — a process often summarized as the central dogma of molecular biology: DNA → RNA → protein.
Core Principles of the Genetic Code
DNA encodes genetic information through the specific sequence of four chemical subunits called nucleotides. Each nucleotide contains a sugar (deoxyribose), a phosphate group, and one of four nitrogenous bases: adenine (A), thymine (T), cytosine (C), and guanine (G). The order of these bases along a DNA strand forms the language of life. Just as the 26 letters of the English alphabet can be arranged to spell millions of different words, the four DNA bases can be arranged in an essentially limitless number of sequences, each carrying different genetic instructions.
Complementary Base Pairing
Triplet Codons
The Central Dogma
Universality of the Code
Degeneracy (Redundancy)
Visual Explanation: DNA Structure and Base Pairing
The diagram illustrates several key features of DNA structure. First, notice that the two strands run in opposite directions — they are antiparallel. The 5′ end of one strand aligns with the 3′ end of the other. Second, the bases always pair in a specific way: A with T, and C with G. This complementary base pairing means that if you know the sequence of one strand, you can predict the sequence of the other. Third, the bases project inward from the backbone, forming the "rungs" of the DNA ladder, while the sugar-phosphate backbone forms the "rails." It is the sequence of these inner bases — not the backbone — that carries genetic information.
From DNA to Protein: The Central Dogma
The information stored in DNA directs the synthesis of proteins through a two-step process. In the first step, called transcription, the enzyme RNA polymerase reads one strand of DNA (the template strand) and builds a complementary strand of messenger RNA (mRNA). RNA uses the base uracil (U) in place of thymine, so wherever the DNA template has an A, the mRNA will have a U. RNA polymerase reads the template strand in the 3′ → 5′ direction and synthesizes the mRNA in the 5′ → 3′ direction.
In the second step, called translation, the ribosome reads the mRNA three bases at a time. Each group of three mRNA bases is a codon. Transfer RNA (tRNA) molecules carry amino acids to the ribosome, matching each codon with its corresponding amino acid through an anticodon — a complementary three-base sequence on the tRNA. As the ribosome moves along the mRNA, amino acids are linked together by peptide bonds, forming a polypeptide chain that folds into a functional protein.
Follow the information flow in the diagram from top to bottom. The DNA template strand is read 3′ → 5′ by RNA polymerase, and a complementary mRNA strand is produced in the 5′ → 3′ direction. Each template base is matched by its RNA complement: T → A, A → U, C → G, and G → C. The resulting mRNA carries the genetic message from the nucleus to the ribosome in the cytoplasm. During translation, the ribosome reads mRNA codons in order from the start codon (AUG) to the first stop codon, linking amino acids into a polypeptide chain. The specific sequence of amino acids determines how the protein folds and what function it performs.
Reading the Genetic Code: The Codon Table
With 64 possible three-base codons and only 20 amino acids, the genetic code is degenerate — meaning that most amino acids are specified by more than one codon. This redundancy is not random; it tends to occur at the third base of the codon, which scientists call the wobble position. The table below shows selected codons and their corresponding amino acids. In practice, scientists use a standard codon chart that lists all 64 possibilities organized by the first, second, and third mRNA bases.
| mRNA Codon | Amino Acid | Role / Notes |
|---|---|---|
| AUG | Methionine (Met) | Start codon — initiates translation |
| UUU, UUC | Phenylalanine (Phe) | Two codons — differ at the wobble position |
| GCU, GCC, GCA, GCG | Alanine (Ala) | Four codons — highly redundant |
| GAG, GAA | Glutamic acid (Glu) | Hydrophilic — important in sickle cell comparison |
| GUG, GUU, GUC, GUA | Valine (Val) | Hydrophobic — replaces Glu in sickle cell hemoglobin |
| UAG, UAA, UGA | — (no amino acid) | Stop codons — signal the ribosome to release the polypeptide |
Notice that a single base change in a codon can change the amino acid or have no effect at all. For example, changing GAG (glutamic acid) to GUG (valine) swaps a hydrophilic amino acid for a hydrophobic one — this is the exact mutation responsible for sickle cell disease. In contrast, changing GCU to GCC still codes for alanine, so the protein is unaffected. This relationship between base sequence and amino acid identity is the foundation of how DNA encodes information.
Worked Example: From DNA to Amino Acid Sequence
Let's work through a complete example of converting a DNA template strand into an amino acid sequence. This exercise practices the skill of tracing the central dogma from start to finish.
When the Code Changes: Types of Mutations
Because protein function depends on the precise sequence of amino acids, any change to the DNA base sequence has the potential to alter the resulting protein. These changes, called mutations, can range from harmless to lethal depending on where they occur and what effect they have on protein structure. Understanding mutation types reinforces how the base sequence encodes information — if it did not matter, mutations would have no consequences.
| Mutation Type | What Happens to DNA | Effect on Protein |
|---|---|---|
| Silent (synonymous) | One base is substituted, but the new codon specifies the same amino acid (e.g., GCU → GCC both code for alanine). | No change — the protein is identical. Redundancy in the genetic code provides a buffer. |
| Missense | One base is substituted, changing the codon to one that specifies a different amino acid (e.g., GAG → GUG changes Glu to Val). | One amino acid changes. May disrupt protein folding and function, or may have little effect depending on location and chemical properties. |
| Nonsense | One base is substituted, creating a premature stop codon (e.g., UAC → UAG changes Tyr codon to Stop). | The protein is truncated (shortened). Often nonfunctional because the full amino acid sequence is needed for proper folding. |
| Frameshift (insertion/deletion) | One or more bases are inserted into or deleted from the sequence (not in multiples of three). | All codons downstream of the mutation are altered. Usually produces a completely nonfunctional protein. |
Beyond the Basics: Gene Expression and Regulation
The central dogma describes the flow of information from DNA to protein, but cells do not simply transcribe all genes all the time. Gene expression is the process by which a specific gene's DNA sequence is read and used to direct protein synthesis. Cells regulate which genes are expressed, when they are expressed, and how much protein is produced. This regulation explains how a single genome can produce hundreds of different cell types — a muscle cell and a neuron contain the same DNA but express very different sets of genes.
| Feature | Basic Central Dogma (This Lesson) | Advanced Gene Expression (Future Topics) |
|---|---|---|
| Scope | DNA → mRNA → protein for a single gene | Regulation of thousands of genes across different tissues and developmental stages |
| mRNA processing | mRNA is a direct copy of the coding sequence | Pre-mRNA is processed: introns are removed, a 5′ cap and poly-A tail are added (eukaryotes) |
| Regulatory elements | Promoter region signals RNA polymerase to begin transcription | Enhancers, silencers, transcription factors, and epigenetic modifications control gene activity |
| Mutation impact | Changes in coding sequence alter amino acids | Mutations in regulatory regions can change when and how much protein is made without altering the protein itself |
As you continue in biology, you will encounter topics like epigenetics, RNA splicing, and gene regulation networks that build on the fundamental concepts covered here. Every one of those advanced topics rests on the principle that the linear sequence of bases in DNA is the primary source of genetic information. Mastering the central dogma gives you the foundation for understanding all of modern molecular biology.
Practice Problems
Summary: How DNA Base Sequences Encode Genetic Information
DNA stores genetic information in the specific sequence of its four nitrogenous bases — adenine, thymine, cytosine, and guanine. The two strands of the double helix are held together by complementary base pairing (A–T and C–G). During transcription, RNA polymerase reads the DNA template strand (3′ → 5′) and produces a complementary mRNA strand (5′ → 3′), substituting uracil for thymine.
During translation, ribosomes read the mRNA in three-base units called codons. Each codon specifies one of 20 amino acids or a stop signal, creating the genetic code (4³ = 64 codons). The code is degenerate — multiple codons can specify the same amino acid, buffering against some mutations. The order of amino acids determines how a protein folds and functions, meaning that the base sequence of DNA ultimately dictates the structure and function of every protein in an organism. This principle — DNA → RNA → Protein — is the central dogma of molecular biology.