# AP Biology Transcription and RNA Processing

> AP Biology · Unit 6: Gene Expression and Regulation
> Source: https://www.owlsprep.com/study/ap-biology-u6-transcription-and-rna-processing/

This module covers core steps of transcription in prokaryotes and eukaryotes, key eukaryotic post-transcriptional modifications, and the role of alternative splicing in generating phenotypic complexity, aligned to the AP Biology CED.

**Prerequisites:** [Central dogma of molecular biology](https://www.owlsprep.com/study/ap-biology-u6-central-dogma/); [Nucleic acid structure and base pairing rules](https://www.owlsprep.com/study/ap-biology-u1-nucleic-acid-structure/); [Prokaryotic vs eukaryotic cell compartmentalization](https://www.owlsprep.com/study/ap-biology-u2-cell-compartmentalization/)

## Learning objectives

- Describe the three core steps of transcription in prokaryotes and eukaryotes
- Explain key post-transcriptional modifications in eukaryotic RNA processing
- Predict mRNA sequences and mature mRNA lengths from given DNA templates
- Connect alternative splicing to phenotypic complexity and gene regulation
- Avoid common exam pitfalls related to transcription and RNA processing

## Core Steps of Transcription

Transcription is the first step of gene expression, where RNA polymerase synthesizes a complementary RNA strand from a DNA template. In addition to messenger RNA (mRNA) that encodes proteins, transcription also produces functional non-coding RNAs like rRNA, tRNA, and siRNA. Transcription occurs in three core steps shared by all organisms, with key differences between prokaryotes and eukaryotes.

**Transcription** — The process of synthesizing a complementary RNA strand from a DNA template, the first step of gene expression in all living organisms.

In initiation, RNA polymerase binds to a promoter sequence upstream of the target gene, assisted by sigma factors (prokaryotes) or general transcription factors and a TATA box (eukaryotes). Only one DNA strand, the *template strand*, is transcribed; the other strand, the *coding strand*, matches the mRNA sequence except thymine replaces uracil. During elongation, RNA polymerase reads the template strand 3'→5', building the mRNA strand 5'→3' by adding complementary ribonucleotides to the free 3' hydroxyl. Unlike DNA polymerase, RNA polymerase does not require a primer to start synthesis. In termination, transcription ends when RNA polymerase reaches a termination sequence: rho-dependent or rho-independent in prokaryotes, and a polyadenylation signal sequence in eukaryotes.

**Worked example:** The 3' → 5' template strand sequence for a short gene is 3' - TAC GAT AGC TTA - 5'. What is the sequence of the matching coding strand, and what is the sequence of the pre-mRNA transcribed from this template?

1. Transcription follows standard base-pairing rules, and strands are always antiparallel, so the coding strand will be complementary to the template and oriented 5' → 3'.
2. The coding strand of DNA matches the mRNA sequence (same 5'→3' orientation) except thymine replaces uracil. Using complementary base pairing, the coding strand sequence is:
3. $$5' - ATG CTA TCG AAT - 3'$$
4. Pre-mRNA is complementary to the template strand, with uracil replacing thymine. This gives the same sequence and orientation as the coding strand, with T swapped for U.
5. The resulting pre-mRNA sequence is:
6. $$5' - AUG CUA UCG AAU - 3'$$

> **tip**
>
> Always check the orientation of the DNA strand given in the question. If given the 5'→3' coding strand, the mRNA sequence is identical except T→U — you do not need to reverse-complement unless given the template strand.

## Eukaryotic Pre-mRNA Processing

RNA processing refers to a set of post-transcriptional modifications that only occur in eukaryotic cells, converting the primary pre-mRNA transcript into mature, translation-ready mRNA that can be exported from the nucleus to the cytoplasm for translation.

**RNA Processing** — Post-transcriptional modifications unique to eukaryotic pre-mRNA that produce mature translation-ready mRNA, including 5' capping, 3' polyadenylation, and RNA splicing.

There are three key modifications to all eukaryotic pre-mRNA: (1) A modified guanine 5' cap added to the 5' end, which protects mRNA from degradation and helps ribosomes bind for translation initiation; (2) A poly-A tail (50-250 adenine nucleotides) added to the 3' end, which also protects against degradation and aids nuclear export; (3) RNA splicing, which removes non-coding intervening sequences called introns and joins coding expressed sequences called exons. Splicing is carried out by the spliceosome, a complex of small nuclear ribonucleoproteins (snRNPs) that recognize splice sites at intron ends.

**Worked example:** A eukaryotic pre-mRNA has the following structure, in order from 5' to 3': 5' untranslated region (part of Exon 1: 50 bp), Exon 1 coding sequence (70 bp), Intron 1 (350 bp), Exon 2 (180 bp), Intron 2 (240 bp), Exon 3 (90 bp), Intron 3 (410 bp), Exon 4 (210 bp, includes 3' untranslated region). After standard (non-alternative) splicing, addition of a 1-nucleotide 5' cap, and a 200-nucleotide poly-A tail, what is the total length of the mature mRNA in nucleotides?

1. Introns are completely removed during splicing, so we only sum the lengths of all exons, including untranslated regions which are retained as part of exons.
2. Sum of all exonic sequence: 50 + 70 + 180 + 90 + 210 = 600 nucleotides.
3. $$50 + 70 + 180 + 90 + 210 = 600$$
4. Add the 5' cap (1 nucleotide) and poly-A tail (200 nucleotides), both of which are part of the final mature mRNA.
5. Total length = 600 + 1 + 200 = 801 nucleotides.
6. $$600 + 1 + 200 = 801$$

> **tip**
>
> Remember that untranslated regions (UTRs) at the 5' and 3' ends of mRNA are part of exons, so they are retained in the mature mRNA, not spliced out like introns.

## Alternative Splicing

Alternative splicing is a regulated post-transcriptional process unique to eukaryotes, where different combinations of exons from the same pre-mRNA are assembled into different mature mRNA molecules. This process explains the 'gene number paradox': humans have only ~20,000 protein-coding genes, but produce hundreds of thousands of distinct proteins. Alternative splicing is also a form of gene regulation, allowing different cell types to produce different protein isoforms with distinct functions from the same gene.

**Alternative Splicing** — A eukaryotic post-transcriptional process that generates multiple distinct mature mRNAs from a single gene by including or excluding different combinations of exons.

**Worked example:** A pre-mRNA has 4 exons in order: Exon 1, Exon 2, Exon 3, Exon 4. Exon 2 can be either included or excluded during splicing, and Exon 3 can be either included or excluded. Exons 1 and 4 are always included, and exon order cannot be rearranged. List all possible unique mature mRNA sequences, and explain why this process increases eukaryotic phenotypic complexity.

1. We retain all combinations of included/excluded Exons 2 and 3, keeping the original order of exons and always including Exons 1 and 4.
2. The four unique mature mRNA sequences are: (1) Exon1-Exon2-Exon3-Exon4, (2) Exon1-Exon2-Exon4, (3) Exon1-Exon3-Exon4, (4) Exon1-Exon4.
3. Each unique mature mRNA has a distinct nucleotide sequence, which is translated into a distinct protein with a different amino acid sequence.
4. A single gene can now produce multiple functional proteins, increasing the total size of the proteome (all proteins an organism can produce) without increasing the number of genes. This allows for greater phenotypic complexity, as different cell types can produce specialized proteins from the same gene.

> **tip**
>
> AP FRQs often ask you to connect alternative splicing to phenotypic variation: remember that one gene → multiple proteins → multiple phenotypes, which explains why organisms with relatively small numbers of genes can have complex traits.

## AP-Style Concept Check

**Check your understanding**

Test your understanding of core concepts with these AP-style questions:

1. A researcher sequences a section of the human genome and identifies the following double-stranded DNA sequence corresponding to the beginning of a gene: 5' - ATGCGTACGTAG... - 3' (Coding Strand), 3' - TACGCATGCATC... - 5' (Template Strand). Which of the following is the correct sequence of the first 10 nucleotides of the pre-mRNA transcribed from this gene?

   - A. 5' - AUGCGUACGU - 3'
   - B. 3' - UACGCAUGCA - 5'
   - C. 5' - ATGCGTACG - 3'
   - D. 5' - UACGCAUGCA - 3'

   *Why:* Correct! The mRNA matches the 5'→3' coding strand sequence, with uracil replacing thymine. Common mistakes include reversing orientation, retaining thymine, or incorrectly complementing the template strand.

2. Beta-thalassemia is a human genetic disorder caused by a point mutation in the 130-nucleotide first intron of the beta-globin pre-mRNA. The mutation prevents the spliceosome from recognizing the intron's 3' splice site, so the entire intron is retained in the mature mRNA. What effect does this mutation have on the resulting beta-globin protein?

   *Why:* This mutation disrupts normal splicing, leading to a frameshift that abolishes protein function, causing the disease phenotype.

## Common pitfalls

- **Wrong:** Stating that RNA processing occurs in prokaryotes.
  - Why it fails: Students confuse prokaryotic coupled transcription-translation with eukaryotic compartmentalization, forgetting processing is only for eukaryotic nuclear pre-mRNA.
  - Correct: Prokaryotic mRNA is not processed; transcription and translation occur simultaneously in the cytoplasm, so no capping, splicing, or polyadenylation occurs.
- **Wrong:** Confusing the template strand and coding strand when writing the mRNA sequence, leading to reversed orientation or incorrect base pairing.
  - Why it fails: Many students memorize that mRNA is complementary to DNA, but forget that only the template strand is transcribed, and the coding strand matches mRNA except T→U.
  - Correct: Always label the orientation of the given DNA strand first: if given 3'→5' template, mRNA is 5'→3' complementary; if given 5'→3' coding, mRNA is same sequence with T→U.
- **Wrong:** Counting introns when calculating the length of mature mRNA.
  - Why it fails: Students remember introns are non-coding, but forget they are completely removed during splicing before the mRNA is mature.
  - Correct: When calculating mature mRNA length, always exclude all intron sequences, sum only exons, then add the 5' cap and poly-A tail length if given.
- **Wrong:** Claiming introns are 'junk DNA' with no function that are just discarded.
  - Why it fails: Many historical sources called introns junk, but AP Biology now tests that introns can have regulatory functions and can be processed into non-coding RNAs.
  - Correct: Never refer to introns as useless junk; acknowledge they are removed from pre-mRNA but many have functional roles in gene regulation.
- **Wrong:** Stating that transcription copies the entire DNA molecule into RNA.
  - Why it fails: Students confuse transcription with DNA replication, which copies the entire genome.
  - Correct: Transcription only copies individual genes (or operons in prokaryotes), initiated at a promoter sequence for the specific gene being expressed.
- **Wrong:** Thinking alternative splicing rearranges the order of exons to make new proteins.
  - Why it fails: Students confuse alternative splicing with DNA recombination.
  - Correct: Alternative splicing only includes or excludes entire exons, never changes the order of exons in the mature mRNA.

## Cheatsheet

| Category | Rule | Notes |
| --- | --- | --- |
| Transcription Direction | RNA synthesized $5' \rightarrow 3'$, reads template $3' \rightarrow 5'$ | Applies to all prokaryotes/eukaryotes |
| mRNA vs DNA Sequence | mRNA matches 5'→3' coding strand: T → U | Complement 3'→5' template if template is given |
| 5' Cap Function | Protects mRNA from degradation, aids ribosome binding | Only added to eukaryotic pre-mRNA |
| Poly-A Tail Function | Protects mRNA from degradation, aids nuclear export | Only added to eukaryotic pre-mRNA |
| RNA Splicing Rule | Introns removed; exons (including UTRs) retained | Spliced by spliceosome made of snRNPs |
| Alternative Splicing Rule | One pre-mRNA → multiple mature mRNAs | Exon order is always preserved, never rearranged |
| Compartmentalization | Transcription/processing in nucleus; translation in cytoplasm (eukaryotes) | Prokaryotes: coupled transcription/translation, no processing |
| RNA Polymerase | Catalyzes phosphodiester bond formation between ribonucleotides | Does not require a primer, unlike DNA polymerase |

## What's next

Transcription and RNA processing form the foundational first step of gene expression, and understanding this process is critical for mastering all subsequent topics in Unit 6. This content sets up the next step of the central dogma: translation, where mature mRNA is decoded to build a functional protein. Without understanding how mature mRNA is produced, you cannot interpret how mutations in splice sites or promoter regions affect protein function, a common topic tested on both AP Biology MCQs and FRQs. Transcription and RNA processing also underpin all mechanisms of gene regulation, the core focus of Unit 6, since many regulatory processes act at the level of transcription initiation or alternative splicing to control when and where proteins are produced. This topic also connects directly to mutations and phenotypic variation, which appear frequently in free response questions.

- [Translation](https://www.owlsprep.com/study/ap-biology-u6-translation/)
- [Mutations and Phenotypic Variation](https://www.owlsprep.com/study/ap-biology-u6-mutations/)
- [Regulation of Gene Expression](https://www.owlsprep.com/study/ap-biology-u6-regulation-of-gene-expression/)

---

From [OwlsPrep](https://www.owlsprep.com) — free study guides for A-Level, IB, AP and IGCSE, written against the official syllabus. Canonical page: https://www.owlsprep.com/study/ap-biology-u6-transcription-and-rna-processing/
