Loading
How cells read the genetic code in DNA and copy it into messenger RNA to direct protein synthesis.
For much of early genetics, scientists understood that genes carry heritable information, but the molecular mechanism by which a gene's instruction is converted into a functional protein remained a profound mystery. The discovery of DNA's double-helix structure in 1953 revealed the architecture of genetic storage, yet the question persisted: how does the information locked within a stable, nuclear DNA molecule reach the ribosomes in the cytoplasm, where proteins are actually assembled? The answer lies in transcription, the process by which DNA is copied into a single-stranded RNA intermediate.
Transcription is the first major step in the Central Dogma of Molecular Biology (DNA → RNA → Protein). Without it, the genetic blueprint would remain an inert archive. Understanding transcription answers the fundamental question: how does a cell convert a static DNA sequence into a dynamic RNA message that can direct the synthesis of a specific protein?
Transcription is the biochemical process by which the nucleotide sequence of one strand of DNA is used as a template to synthesize a complementary strand of RNA. This process takes place in the nucleus of eukaryotic cells (or the nucleoid region of prokaryotic cells) and is catalyzed by the enzyme RNA polymerase. Several foundational principles govern how transcription operates.
The diagram below illustrates the transcription process, showing how RNA polymerase opens the DNA double helix at a transcription bubble, reads the template strand, and synthesizes a complementary mRNA molecule. Notice the directionality: the template strand is read 3′ → 5′ while the mRNA is built 5′ → 3′.
As RNA polymerase moves along the DNA, it unwinds a short region of roughly 12–17 base pairs called the transcription bubble. Within this bubble, the template strand is transiently exposed and available for base pairing with incoming ribonucleoside triphosphates (NTPs). Behind the polymerase, the DNA strands re-anneal, and the newly synthesized mRNA peels away as a single-stranded molecule. The resulting mRNA carries the same sequence as the coding strand (with uracil in place of thymine) and will ultimately be read by ribosomes during translation.
Transcription proceeds through three distinct stages: initiation, elongation, and termination. Although the fundamental chemistry is the same in prokaryotes and eukaryotes, the molecular details differ. Below we describe the general mechanism, emphasizing the eukaryotic process where relevant.
Initiation is the most highly regulated step and determines which genes are transcribed and when. In prokaryotes, the sigma (σ) factor associates with the core RNA polymerase to form the holoenzyme, which recognizes conserved promoter sequences (the −10 region, also called the Pribnow box, and the −35 region). In eukaryotes, the process is more elaborate: general transcription factors (TFIIA, TFIIB, TFIID, TFIIE, TFIIF, and TFIIH) assemble at the promoter to form the pre-initiation complex (PIC). The key promoter element in many eukaryotic genes is the TATA box, located approximately 25–30 base pairs upstream of the transcription start site, which is recognized by the TATA-binding protein (TBP), a subunit of TFIID.
Once the pre-initiation complex is formed, TFIIH (which has helicase and kinase activities) unwinds the DNA to create the transcription bubble and phosphorylates the C-terminal domain (CTD) of RNA polymerase II at serine-5 residues. This phosphorylation event triggers promoter clearance — the polymerase escapes the promoter and enters the elongation phase.
During elongation, RNA polymerase moves along the template strand at a rate of approximately 40–80 nucleotides per second in eukaryotes (faster in prokaryotes). The enzyme maintains a transcription bubble of about 12–17 base pairs. At the active site, incoming NTPs are selected by complementary base pairing with the template and are joined to the 3′-OH end of the growing mRNA via a phosphodiester bond. The energy for bond formation comes from the hydrolysis of the high-energy phosphate bonds of the NTP substrate. Behind the polymerase, an RNA–DNA hybrid of about 8 base pairs exists transiently before the RNA strand peels off.
In eukaryotes, as RNA polymerase II elongates, additional phosphorylation of the CTD at serine-2 recruits factors for RNA processing: the 5′ cap is added when the transcript is about 20–30 nucleotides long, and splicing factors begin to assemble on the pre-mRNA co-transcriptionally.
Termination signals the end of transcription and the release of both the mRNA and RNA polymerase from the DNA template. In prokaryotes, two mechanisms are common: Rho-independent (intrinsic) termination, where a GC-rich hairpin loop followed by a poly-U tract in the RNA destabilizes the elongation complex, and Rho-dependent termination, where the Rho helicase protein chases the polymerase and unwinds the RNA–DNA hybrid. In eukaryotes, termination of RNA Pol II transcripts is coupled to cleavage and polyadenylation: the pre-mRNA is cleaved at a site downstream of the AAUAAA polyadenylation signal, a poly(A) tail is added by poly(A) polymerase, and the polymerase eventually dissociates.
The following flowchart summarizes the major molecular events of transcription in eukaryotes, from promoter recognition through mRNA release. Each box represents a key molecular event, with the participating factors noted.
This flowchart emphasizes that transcription in eukaryotes is tightly coupled with RNA processing events. The 5′ cap (a 7-methylguanosine linked by a 5′-to-5′ triphosphate bridge) is added early during elongation, splicing of introns occurs co-transcriptionally as the pre-mRNA emerges from the polymerase, and the 3′ poly(A) tail is added during termination. These processing steps convert the pre-mRNA (primary transcript) into a mature mRNA ready for export and translation.
| Feature | Prokaryotic Transcription | Eukaryotic Transcription |
|---|---|---|
| RNA Polymerase | Single core enzyme (α₂ββ′ω) + σ factor | Three main types: Pol I (rRNA), Pol II (mRNA), Pol III (tRNA, 5S rRNA) |
| Promoter Elements | −10 (Pribnow box) and −35 regions | TATA box (~−25), Inr, DPE, enhancers/silencers |
| Transcription Factors | σ factor (initiation only) | TFIIA, B, D, E, F, H + Mediator complex |
| mRNA Processing | None — mRNA is translated co-transcriptionally | 5′ capping, splicing, 3′ polyadenylation |
| Termination | Rho-dependent or Rho-independent (hairpin) | Coupled to cleavage/polyadenylation (Pol II) |
| Coupling to Translation | Yes — ribosomes attach to mRNA during transcription | No — mRNA must be exported to cytoplasm first |
The following problem walks through how to determine the mRNA sequence produced from a given DNA template strand and calculate the time required for transcription of a gene.
3′ — T A C G G A T T C A A G C T A — 5′Template 3′: T A C G G A T T C A A G C T A :5′mRNA 5′: A U G C C U A A G U U C G A U :3′Transcription and DNA replication are both template-directed nucleic acid synthesis processes, but they differ in fundamental ways. Understanding these differences highlights why cells use two separate molecular machines for two distinct purposes.
| Feature | Transcription | DNA Replication |
|---|---|---|
| Product | Single-stranded RNA | Double-stranded DNA |
| Enzyme | RNA polymerase | DNA polymerase (+ helicase, primase, ligase, etc.) |
| Template | One strand of DNA (for a given gene) | Both strands of DNA |
| Primer needed? | No — initiates de novo | Yes — requires an RNA primer |
| Nucleotides used | NTPs (ATP, UTP, GTP, CTP) | dNTPs (dATP, dTTP, dGTP, dCTP) |
| Sugar in product | Ribose (2′-OH present) | Deoxyribose (2′-H) |
| Proofreading | Limited (error rate ~10⁻⁵) | Extensive 3′→5′ exonuclease (error rate ~10⁻⁹) |
| Scope | Selected genes, as needed | Entire genome, once per cell cycle |
| Direction of synthesis | 5′ → 3′ | 5′ → 3′ |
The basic mechanism of transcription described above is a foundation for understanding a vast landscape of regulatory and biotechnological topics. As you advance in molecular biology, you will encounter increasingly sophisticated layers of transcriptional control and applications.
Gene expression is regulated at the transcriptional level by transcription factors — proteins that bind to enhancers, silencers, and other regulatory DNA elements to activate or repress transcription. Activators recruit RNA polymerase and coactivators (such as the Mediator complex), while repressors block polymerase binding or recruit chromatin-remodeling complexes that condense chromatin. Beyond sequence-specific transcription factors, epigenetic modifications — including histone acetylation, histone methylation, and DNA methylation — alter chromatin accessibility without changing the DNA sequence. These modifications create a layer of heritable gene regulation that governs cell identity and development.
Not all transcription products encode proteins. Cells produce a diverse repertoire of non-coding RNAs (ncRNAs) including ribosomal RNA (rRNA), transfer RNA (tRNA), microRNAs (miRNAs), long non-coding RNAs (lncRNAs), and small interfering RNAs (siRNAs). Many of these are transcribed by RNA Pol I or Pol III, and they play essential roles in translation, RNA processing, gene silencing, and chromatin organization.
Mutations in promoters, transcription factors, or the transcription machinery itself can cause disease. For example, mutations in the BRCA1 gene disrupt transcription-coupled DNA repair and increase breast cancer risk. Drugs targeting transcription, such as α-amanitin (which inhibits RNA Pol II) and rifampicin (which inhibits bacterial RNA polymerase), are invaluable in research and medicine. In biotechnology, understanding promoters allows scientists to engineer gene expression systems — placing a gene under a strong, inducible promoter to produce recombinant proteins at will.
| Concept | Basic Transcription | Advanced Extension |
|---|---|---|
| Promoter recognition | TATA box + general TFs | Enhancers, silencers, insulators, Mediator complex, 3D chromatin looping |
| RNA processing | Capping, splicing, poly(A) | Alternative splicing, RNA editing, RNA surveillance (NMD) |
| Gene regulation | Activators / repressors | Epigenetics, chromatin remodeling, phase separation of transcription hubs |
| Transcription errors | Limited proofreading | Transcription-coupled nucleotide excision repair (TC-NER) |
As you progress, you will see that transcription is not merely a mechanical copying process but rather a deeply regulated, dynamic, and context-dependent event at the heart of cellular decision-making.
Test your understanding with these five problems of increasing difficulty.
5′-ATGCCCGAATTCGGA-3′, write: (a) the template strand, and (b) the mRNA sequence that would be transcribed.Transcription is the first step of gene expression, in which the enzyme RNA polymerase reads the template (antisense) strand of DNA in the 3′ → 5′ direction and synthesizes a complementary messenger RNA (mRNA) molecule in the 5′ → 3′ direction. The process unfolds in three stages: initiation, where RNA polymerase binds to the promoter (aided by sigma factors in prokaryotes or general transcription factors and the TATA box in eukaryotes); elongation, where the polymerase moves along the template, adding ribonucleotides (with uracil replacing thymine); and termination, where transcription stops and the mRNA is released. RNA uses ribose sugar and does not require a primer, distinguishing transcription from DNA replication.
In eukaryotes, the primary transcript (pre-mRNA) undergoes three critical processing steps: addition of a 5′ cap, removal of introns by splicing, and addition of a 3′ poly(A) tail. These modifications protect the mRNA, facilitate nuclear export, and regulate translational efficiency. Transcription is tightly regulated by transcription factors, enhancers and silencers, and epigenetic modifications, making it a central control point for determining which genes are expressed in a given cell type or condition. Understanding transcription provides the essential foundation for studying gene regulation, biotechnology, and the molecular basis of disease.
Keep learning with more lessons from the same subject.