7.4. Biopolymer Structure of Nucleic Acids

The polymeric backbone of nucleic acids is repeating units of phosphate-saccharide (Figure 7.14). The nucleobase components are not part of the polymer backbone, they “branch” off in the same direction together.

image

Figure 7.14 – Simplified Biopolymeric Structure of a Nucleic Acid.

While the nucleobase branches are all planar/flat, the saccharide-phosphate biopolymer is not. As a result, the actual three-dimensional shape of nucleic acids is complex. Even in the absence of other effects the stereochemistry of the saccharides causes the chain to twist slowly, helping to contribute to the overall shape of the biopolymer (Figure 7.15). In most nucleic acid polymers this is aided by pi-stacking (π-stacking) between sequential nucleobases on the biopolymer.

image

Figure 7.15 – Approximated Three-Dimensional Shape of the Biopolymeric Structure of a Nucleic Acid.

7.4.1. Nucleic Acid Terminology: 5’-Terminus and 3’-Terminus

Nucleic acids are typically made from a combination of all of the possible nucleotide monomers. Like peptides, often the exact sequence (the order that they are connected in) is important. To avoid ambiguity a naming convention for nucleic acid chains was established.

The end of the chain that would contain the last alcohol at position 5 of a saccharide is called the 5’-terminus (spoken aloud as the “five prime terminus”; Figure 7.16). The sequence is read starting from this nucleotide. The end of the chain that would contain the last alcohol at position 3 of a saccharide is called the 3’-terminus (spoken aloud as the “three prime terminus”). The sequence is read ending at this nucleotide. The orientation and/or perspective of the drawing is irrelevant, the chain is always read from the 5’-terminus to the 3’-terminus.

image

Figure 7.16 – Generalized Examples of Reading Nucleic Acid Sequences from 5’-Terminus to 3’-Terminus Regardless of Drawing Orientation/Perspective.

In both cases the “prime” (’) component of the term comes from applying specific numbering rules for naming nucleosides (Figure 7.17). Under the formal rules the nucleobase positions are numbered first, then the saccharide positions are numbered and given “prime” symbols to differentiate them from the positions of the nucleobase. It is not necessary to be able to number and formally name these compounds, this is provided only for context.

image

Figure 7.17 – Formal IUPAC Numbering of Adenosine.

In most nucleic acids both termini are alcohols. However, as with peptides the 5’-terminus does not have to be the parent primary alcohol functional group; the 3’-terminus does not have to be the parent secondary alcohol functional group (Figure 7.18). Nucleic acid chains that have been modified and/or that are in a different protonation state are still read from the 5’-terminus to the 3’-terminus. The functional group itself may be modified, including to phosphates or other functional groups (e.g. ethers, esters, etc.).

image

Figure 7.18 – Generalized Examples of Reading Nucleic Acid Sequences from 5’-Terminus to 3’-Terminus Regardless of Functional Group Modification.

7.4.2. Nucleic Acid Nomenclature: Describing Nucleic Acid Chains

Under the IUPAC rules the nucleic acid sequence may be expressed in several different ways. However, a simple method is commonly used (Figure 7.19). This expresses the sequence of nucleobases in order from 5’-terminus to 3’-terminus using the one-letter codes for the nucleobases. Technically the use of three-letter codes is also valid but not recommended; the use of three-letter codes to describe sequences is discouraged but may be encountered in some sources. Individual nucleobases may or may not be separated by hyphens when using the one-letter codes. The same system is used to describe the sequence regardless of which class of nucleic acid is being described (DNA or RNA).

image

Figure 7.19 – Examples of Naming Nucleic Acid Sequences for Two Trimers.

There are many other one-letter codes often included in nucleic acid sequences. These are used to represent uncommon nucleobases or ambiguity. For example, there is a letter code to represent “inosine”, another to represent “a purine nucleobase”, and another to represent “any nucleobase other than A”. This text will avoid using these letter codes but they may be encountered in other sources.

Technically, if the 5’ or 3’ positions are modified to different functional groups there are ways to describe the change(s) made. However, at an introductory level it is often difficult to know when/how to describe modifications. This text will avoid examples using these variations but they may be encountered in other sources.

7.4.3. The Double Helix and Complexity

Although not the focus of this text, it is important to recognize that nucleic acids are a very complex class of compounds that can perform a wide array of functions. For example, there are several different “kinds” or ribonucleic acids (e.g. mRNA, rRNA, tRNA, etc.) that perform significantly different functions in the cell.

At the same time, the actual structures adopted by deoxyribonucleic acids are incredibly complex (Figure 7.20). Once generated they typically combine with a complimentary strand made from the base pairs that match the sequence. Efficient hydrogen bonding holds the two polymers together. The structure then spirals around itself as the sequence extends due to a combination of stereochemistry/three-dimensional shape (the saccharides) and pi-stacking interactions between sequential nucleobases. This is the so-called “double helix”.

image

Figure 7.20 – Approximated Three-Dimensional Shape and Ribbon Model of the Double Helix Structure of a Nucleic Acid Dimer with B-DNA Helical Shape.

Although the classical image of a double helix already appears relatively complex, it is in fact significantly more so: there are a large number of prominent three-dimensional traits (e.g. major vs. minor grooves), different helix shapes (e.g. A-DNA vs. B-DNA vs. Z-DNA), folding and quaternary structures, chemical modifications, etc. that all serve different roles in how the DNA functions in the cell.

Most institutions have multiple entire courses dedicated to understanding the structure and function of deoxyribonucleic acid and/or ribonucleic acid. At an introductory chemistry level only a very simplified overview of the basics is discussed. This text will focus on the simple chemical differences between the two types of nucleic acids, why they are different, and how to synthesize simple nucleic acid polymers. Discussion of the complex three-dimensional shapes, their functions, how cells interpret sequences of nucleobases, and in-depth discussion of the enzymes involved is left for advanced courses.