Supporting material

Documents and resources

The sources behind each Aequorea video: scientific papers, Nobel lectures, open textbook chapters and our own maps, organised by video and topic. Every document carries its full citation; those hosted here show their licence, and third-party ones link to their original publication.

All
All
By Aequorea
Third-party
Document text

Everything here describes transcription in prokaryotes, specifically Escherichia coli . Eukaryotes follow the same principle, but they use three different RNA polymerases instead of one, they need additional transcription factors, and they process the RNA before using it. Where no organism is named, assume prokaryote.

Before we start

The five phases of transcription, one per slide: the promoter, the bubble, initiation, elongation and termination. Plus the molecule the process produces.

In this document

How a cell reads one gene and leaves the rest of the genome untouched. The whole process in prokaryotes, with E. coli as the model. From DNA to message

Central dogma

Switching genes on and off is the basis of cell differentiation, of responding to the environment, and of metabolic adaptation. Without transcriptional regulation there would be no tissues and no way to react to a change in nutrients.

Why it matters

E. coli carries around 4,400 genes in its 4.6 million base pairs, and under normal growth conditions it expresses about half of them. A human cell has roughly 20,000 coding genes, and each cell type uses a different fraction. A neuron and a liver cell carry the same genome; what sets them apart is which part of it they transcribe.

The real scale

Replication copies the entire genome, once per cell division, and produces DNA. Transcription copies specific fragments, many times over, at any point in the cycle, and produces RNA. One gene can be transcribed hundreds of times while the one beside it never becomes a message at all.

What changes from replication

At any given moment a cell transcribes only some of them. RNA polymerase has to find where to start . 4,400Copy only what’s needed genes in E. coli

The problem

RNA polymerase makes roughly one error per 10⁴ nucleotides. DNA polymerase III, with its proofreading systems, gets down to one per 10⁹. The difference tracks how long the product lasts: a faulty RNA degrades in minutes, while an error in DNA is inherited by every daughter cell.

The precision

Only one of the two strands gets transcribed. That one is the template strand, also called antisense. The other is the coding or sense strand, because its sequence matches the resulting RNA except that it carries thymine where RNA carries uracil. Which of the two serves as template depends on the gene and changes along the chromosome.

The template strand

RNA polymerase can start a chain from nothing. DNA polymerases can only extend an existing chain, which is why replication needs primase to synthesize a stretch of RNA before DNA polymerase III can bind. Primase is itself a specialized RNA polymerase.

Why no primer is needed

Copies one gene. Starts without a primer . Uses one strand as template. Produces RNA. Copies the entire genome. Needs a primer. Uses both strands as template. Produces DNA. Replication T ranscriptionTwo ways to copy

Comparison

A bacterium carries several sigma factors, each specialized in a different set of promoters. Switching sigma factors activates an entire genetic program. E. coli uses a different one for its heat shock response, and in Bacillus subtilis a cascade of sigma factors drives sporulation.

The other sigma factors

An adenine-thymine pair holds together with only two hydrogen bonds, against the three of guanine-cytosine. That makes the −10 box the easiest stretch of the promoter to open, and it is exactly where the transcription bubble forms in the next step.

T rich

The RNA polymerase of E. coli has a core enzyme that synthesizes RNA but does not recognize where to begin. The sigma factor supplies that specificity: bound to the core it forms the holoenzyme , and the holoenzyme recognizes the promoter’s two consensus sequences. The numbers −35 and −10 give their position relative to the start site, numbered +1 by convention.

The mechanism

The genome has no index. The sigma factor recognizes two sequences that mark where a gene begins: the −35 box and the −10 box , also known as the Pribnow box. Sigma factorFinding where to start

Step 1 of 5

The bubble travels along with RNA polymerase. DNA opens ahead of it and winds back up behind, so only a tiny fraction of the chromosome is ever exposed. That allows transcription without destabilizing the genome.

The moving bubble

Opening the helix in replication takes a dedicated helicase that burns ATP with every cycle. Here RNA polymerase opens, reads and closes on its own. The difference lies in the job itself: opening 15 base pairs at a time is a different task from opening an entire genome.

The comparison with replication

In the closed complex the holoenzyme sits on the DNA and the double helix stays intact. The shift to an open complex separates the strands locally, starting at the −10 box. The sigma factor drives it without spending ATP, since the energy comes from binding to the DNA itself. The bubble measures 12 to 15 base pairs and keeps that size throughout.

The mechanism

Binding to the promoter forms a closed complex . Once the strands separate it becomes an open complex : a transcription bubble of 12 to 15 base pairs. HoloenzymeOpening the bubble

Step 2 of 5

After about ten nucleotides the sigma factor detaches. Its job was to locate the promoter and open the helix, and it takes no further part. The core enzyme carries on alone, and the sigma factor is free to bind another core and find another promoter.

The sigma factor leaves

Before committing, RNA polymerase synthesizes transcripts of two to ten nucleotides and releases them, repeating the attempt several times. Only once it clears that critical length does it break free of the promoter and settle into stable elongation.

Abortive initiation

DNA polymerases only extend an existing chain. RNA polymerases start one. That structural difference explains why replication needs primers and transcription does not. The first nucleotide of the transcript sits opposite the +1 site and keeps all three of its phosphates, which no other nucleotide in the chain does.

The mechanism

With the +1 site exposed, RNA polymerase places the first ribonucleotide. It needs no primer : it can start the chain from nothing. RNA polymeraseInitiation

Step 3 of 5

RNA polymerase backs up when it adds the wrong nucleotide, cuts out the faulty stretch and tries again. The mechanism exists, but it is far less thorough than that of DNA polymerase III, and the final error rate stays at one per 10⁴ nucleotides.

Proofreading

Cytosine deaminates spontaneously and turns into uracil. If DNA carried uracil as a normal base, repair systems could not tell a legitimate uracil from a damaged cytosine. Thymine, which is uracil with a methyl group, solves the problem: any uracil that shows up in DNA is an error, and uracil-DNA glycosylase removes it. RNA lasts minutes and needs no such protection, so it keeps uracil.

Why uracil

RNA polymerase moves at around 50 nucleotides per second, twenty times slower than DNA polymerase III. Positive supercoils build up ahead of the bubble and negative ones behind it. Gyrase clears the positive ones, as it does in replication, and topoisomerase I relaxes the negative ones.

The mechanism

It reads the template 3′→5′ and builds the RNA 5′→3′, adding adenine, uracil , cytosine and guanine. Up ahead, gyrase releases the torsion. Core enzymeReading and writing

Step 4 of 5

Both mechanisms coexist in the same organism, and each gene uses whichever its sequence dictates. The intrinsic one consumes no energy. The rho-dependent one adds a control point, since the cell can regulate how much rho factor is available and, with it, where transcription stops.

Why there are two ways

The rho factor is an ATPase with helicase activity. It binds the rut site, from rho utilization, as soon as it appears on the transcript, and travels along the RNA 5 ′→3′ faster than RNA polymerase. On reaching it, rho unwinds the RNA-DNA hybrid and RNA polymerase comes off the template.

Assisted mechanism

The guanine-cytosine hairpin pulls the RNA backward, and directly below it sits the run of uracils paired with adenines on the template. A uracil-adenine pair holds with two hydrogen bonds and forms the least stable RNA-DNA hybrid possible. Tension from the hairpin over a weak pairing is enough to release the transcript.

Intrinsic mechanism

The rho factor binds a rut sequence on the transcript, moves along burning ATP until it reaches RNA polymerase, and knocks it loose. Rho-dependent The RNA folds into a G-C rich hairpin , followed by a run of uracils . The hybrid destabilizes with no protein involved. Rho-independentRho factor T ermination

Step 5 of 5

A bacterial messenger RNA has a half-life of a few minutes. That short life lets a cell stop producing a protein almost immediately, just by ceasing to transcribe its gene. With a stable messenger, switching a gene off would have no effect for hours.

How long it lasts

Prokaryotes have no nucleus, so messenger RNA shares a compartment with ribosomes from the instant it appears. Translation begins on a transcript that is still being synthesized: one RNA can have RNA polymerase writing at one end and several ribosomes reading at the other. In eukaryotes the nuclear membrane separates the two processes, and that coupling does not happen.

Coupled transcription

A single-stranded messenger RNA, complementary to the template strand and identical to the coding strand except for uracil. Its sequence already holds, in groups of three bases, the instructions for assembling a protein. How those triplets are read is the subject of folder 4.

What it is

A single strand of RNAT

He result

Complementary to the template and ready to be translated. In prokaryotes, ribosomes start reading it before it’s finished being written.

- Alberts, B., Heald, R., Johnson, A., Morgan, D., Raff, M., Roberts, K., & Walter, P. (2022). Molecular biology of the cell (7th ed.). W. W. Norton & Company. - Feklístov, A., Sharon, B. D., Darst, S. A., & Gross, C. A. (2014). Bacterial sigma factors: A historical, structural, and genomic perspective. Annual Review of Microbiology, 68, 357–376. https://doi.org/10.1146/annurev-micro-092412-155737 - Nelson, D. L., Cox, M. M., & Hoskins, A. A. (2021). Lehninger principles of biochemistry (8th ed.). W. H. Freeman. - Ray-Soni, A., Bellecourt, M. J., & Landick, R. (2016). Mechanisms of bacterial transcription termination: All good things must end. Annual Review of Biochemistry, 85, 319–347. https://doi.org/10.1146/annurev-biochem-060815-014844 - Watson, J. D., Baker, T. A., Bell, S. P., Gann, A., Levine, M., & Losick, R. (2014). Molecular biology of the gene (7th ed.). Pearson.

References

The interactive Central Dogma minigame is at aequorea.net, alongside the rest of the series’ downloadable resources. The full video walks through every one of these steps with 3D animation.

Keep studying

The messenger RNA carries the complete instruction, and a machine is needed to read it and turn those bases into amino acids. That is the job of the ribosome and of translation, folder 4 in this series.

S next

The whole process, animated and explained in full, in the Central Dogma video. The messageis written