DNA and RNA are polymers, and like the other biomolecules in this chapter they are built from chemistry you already have: an acetal linking a sugar to a base, a phosphate ester joining one unit to the next, and hydrogen bonds holding two chains together. The information content is biology's; the bonds are this course's.
Three pieces, two names
Each repeating unit is built from a sugar, a nitrogenous base and a phosphate, and the two names for partial assemblies are worth keeping straight because questions turn on them:
- Nucleoside = sugar + base.
- Nucleotide = sugar + base + phosphate. The extra t is the phosphate.
The sugar is the difference between the two polymers. RNA uses ribose; DNA uses 2-deoxyribose, which is ribose missing the OH at C2 — which is exactly what "deoxy" is saying.
The bases, in two families
- Purines — two fused rings: adenine (A) and guanine (G).
- Pyrimidines — one ring: cytosine (C), thymine (T) in DNA and uracil (U) in RNA.
Thymine and uracil differ by a single methyl group, which thymine has and uracil does not. The bases attach to the sugar's anomeric carbon through nitrogen, forming an N-glycosidic bond — the same acetal-type linkage as in the disaccharides earlier in this chapter, with nitrogen in place of oxygen.
The backbone is a chain of phosphate diesters
Units are joined by a phosphodiester bond: a phosphate esterified to the 3′ carbon of one sugar and the 5′ carbon of the next. The result is an alternating sugar–phosphate backbone with the bases hanging off it.
That backbone carries one negative charge per phosphate at physiological pH, which makes nucleic acids strongly polyanionic — the reason they migrate in an electric field and the basis of gel electrophoresis.
Because the two ends of the chain are different — one has a free 5′ phosphate and the other a free 3′ OH — a sequence has a direction, and it is written 5′ to 3′ by convention. As with a peptide's N-to-C convention, writing the sequence backwards describes a different molecule.
Base pairing: the part that is pure hydrogen bonding
Two strands associate because the bases pair specifically, and the specificity is entirely a matter of which hydrogen-bond donors and acceptors line up:
| Pair | Hydrogen bonds | Note |
|---|---|---|
| A–T (A–U in RNA) | 2 | Purine with pyrimidine |
| G–C | 3 | Purine with pyrimidine; the stronger pair |
Two structural consequences follow directly. Every pair is a purine with a pyrimidine, so each rung of the ladder is the same width and the helix has a constant diameter — two purines would be too wide and two pyrimidines too narrow. And because G–C has three hydrogen bonds to A–T's two, a GC-rich stretch of DNA takes more energy to separate, which is why GC content predicts melting temperature.
The two strands run in opposite directions — antiparallel, one 5′→3′ alongside the other 3′→5′ — which is what allows the bases to face each other correctly.
One DNA strand reads 5′-ATGC-3′. Give the complementary strand.
Pair each base: A with T, T with A, G with C, C with G, giving TACG.
Get the direction right. The strands are antiparallel, so the complement read alongside the original runs 3′→5′. Written conventionally, 5′→3′, it is reversed: 5′-GCAT-3′.
The reversal is where these questions are lost. Pairing the bases is mechanical; remembering that the answer has to be written in the other direction is the actual content.
Why hydrogen bonds are the right choice
It is worth asking why the two strands are held by hydrogen bonds rather than something stronger. An individual hydrogen bond is weak — a few percent of a covalent bond — and easily broken at biological temperatures.
That is the point. The strands have to separate for the molecule to be read or copied, and then reassociate. A covalent link would make the information permanent and unusable. Thousands of weak interactions give a structure that is stable overall and can still be opened locally, which is precisely what is required — and it is the same "weak individually, decisive in bulk" argument that explains protein folding in the previous section.
What carries forward
This is the last section of the biomolecules chapter, and a good summary of what the chapter was for. A nucleotide is an acetal, an ester and an amine-bearing heterocycle; a double helix is hydrogen bonding and shape complementarity. None of it needs new chemistry — it needs the chemistry you have, applied to a molecule that happens to be alive.