7. What DNA Does (1,139;7/31)
- lscole
- Apr 15, 2025
- 5 min read
Updated: Jul 31
How does a string of letters--really, a string of molecules--actually do anything? In this chapter, we shift our focus from structure to function: how DNA works.
The genetic code
If we were to walk down a stretch of DNA we could announce the letters on the strand as we passed: "ATGTCGGATAGATGA", for example. A code is contained in these 15 letters. Every protein in your body--from enzymes to structural proteins--starts as a DNA sequence like this, but longer.
A primary purpose of the genetic code is to make proteins. But DNA does so indirectly using a certain type of RNA-- mRNA--as an intermediary (the "m" stands for "messenger"). Thus, the DNA code will effectively be copied into a kind of mRNA photocopy that will in turn be used to synthesize the protein. This process will be covered in the next chapter.
For now, two things matter about RNA. First, like DNA, RNA is a chain of nucleotides--but with slightly different nucleotide chemistry. DNA uses DNA nucleotides while RNA uses RNA nucleotides. Their respective backbone components differ only by one very small chemical attachment.
Second, one of the four bases differs between DNA nucleotides and RNA nucleotides. DNA uses thymine (T) base while RNA uses a uracil (U) base. So, if a stretch of DNA reads ATGTAAC, the corresponding RNA sequence would read AUGUAAC.
Codons
To make a protein, the cell has to know the order of its amino acids. Fortunately, the order of the nucleotides in a given protein's gene and its mRNA photocopy convey the order.
The cell's instructions for making proteins uses a code that assigns every possible three-letter combination of nucleotide bases (called codons) to specific amino acids. This is the genetic code.

Returning to our example, how would we interpret the sequence ATGTCGGATAGATGA?
We would interpret it by grouping. That same sequence can also be viewed as a stretch of three letter codons. Grouped in threes, the sequence becomes: ATG TCG GAT AGA TGA.
Think of a three-letter codon as a word in a spoken language. Think of a gene as a sentence. A gene (sentence) is a string of codons (words) that has meaning.
For example, if we have a piece of DNA 450 bases long, that would equate to 150 codons (450 bases divided by three bases per codon). And with a 150 codon gene, the cell can make a protein that's 150 amino acids long. Just as DNA’s structure makes copying possible, its nucleotide sequence makes it readable.
When this DNA sequence is turned into an mRNA molecule, the "T" (thymine) nucleotide in DNA will be replaced with a "U" (uracil) nucleotide in the mRNA. So the gene's codons in the language of mRNA would be: AUG UCG GAU AGA UGA.
These triplets aren't just letters--they're instructions. Each codon corresponds to a specific amino acid. The first codon in our gene is ATG. That codon specifies the amino acid methionine (Met in the chart). The first amino acid in this protein is methionine.
The second amino acid in our protein is specified by the codon UCG. UCG corresponds to the amino acid serine (Ser). Now we know that the first two amino acids in our protein are Met-Ser.
How Many Codons Are Enough?
At this point, something might be bothering us. Given there are four different nucleotides and three code positions, there are four to the third power--that is, 64--possible codons.
There are many more codons available to the cell (64) than there are amino acids (20). In theory, we have too many codons. This seems wasteful.
But instead of waste, this turns out to be a form of built-in protection. The cell deals with this using what's referred to as redundancy, meaning most amino acids are identified by more than one codon. For example, the amino acid alanine (Ala) is associated with four codons: GCG, GCA, GCC and GCU.
Notice that the difference between these four codons is the last nucleotide. If you scan the codon table, you'll see that the first two nucleotides of a codon seem to dominate. Much of the code’s resilience comes from the third position in each codon, where changes often have no effect at all.
But even when mutations alter other positions, the code is structured so that similar codons tend to specify chemically similar amino acids. Together, these features make the system surprisingly tolerant to error.
Identifying a gene
So far, we’ve treated DNA as if it were one long continuous message. But in reality, only certain stretches--genes--are used to make proteins. It might surprise you that 98.5% of the genome does not directly code for proteins.
So how does the cell identify genes contained in long stretches of nucleotides There’s no single signal. Instead, multiple local features combine to make genes recognizable.
For example, genes are often preceded by CpG islands--long stretches of DNA enriched with CG dinucleotides. Genes also tend to be in loosely packed DNA called euchromatin with specific chemical markings rather than in tightly packed heterochromatin lacking the markings.
There are also sequences both within a gene and just in front that play key roles. Near the start of genes--that is, near the transcription start site and also often near these CpG islands--are promoter sequences. Promoter sequences are recognized by the transcription machinery--that is, the group of enzymes that turn DNA genes into mRNAs--which assembles nearby. This occurs just prior to transcription.
The core promoter sequence--the minimal region needed to effect transcription--is usually small. It's on the order of a few dozen base pairs. But it typically sits inside a much longer stretch of DNA, often a few hundred base pairs long, that helps regulate when and how strongly a gene is expressed.
In addition, certain codons identify where protein coding begins. The protein-coding portion of most genes begins with a start codon, usually ATG (AUG in RNA-speak). This start codon is indicated in green font in the codon table.
The start codon also codes for the amino acid methionine (Met). So the first amino acid in many proteins is methionine, although it's sometimes removed.
Genes also have stop codons that identify their termini. There are three stop codons: UAA, UAG, and UGA. When the enzymes that make proteins encounter a stop codon on an mRNA, protein synthesis ends.
So far, we’ve focused on the code itself--how DNA stores instructions for building proteins. But remember, there is no central reader and no plan--just molecules crashing into each other and interacting locally. And yet, the chemical code on mRNAs is reliably turned into proteins.
In the next chapter, we’ll follow the flow of information from DNA to RNA to protein--this idea was first put forth in 1957 by Francis Crick of Watson and Crick fame. It's important enough that he got away with calling it the central dogma of molecular biology.

Comments