Skip to content

Proteins

By the end of this section, you will be able to:

  • Describe the fundamental structure of an amino acid
  • Describe the chemical structures of proteins
  • Summarize the unique characteristics of proteins

At the beginning of this chapter, a famous experiment was described in which scientists synthesized amino acids under conditions simulating those present on earth long before the evolution of life as we know it. These compounds are capable of bonding together in essentially any number, yielding molecules of essentially any size that possess a wide array of physical and chemical properties and perform numerous functions vital to all organisms. The molecules derived from amino acids can function as structural components of cells and subcellular entities, as sources of nutrients, as atom- and energy-storage reservoirs, and as functional species such as hormones, enzymes, receptors, and transport molecules.

Amino Acids and Peptide Bonds

An amino acid is an organic molecule in which a hydrogen atom, a carboxyl group (–COOH), and an amino group (–NH₂) are all bonded to the same carbon atom, the so-called α carbon. The fourth group bonded to the α carbon varies among the different amino acids and is called a residue or a side chain, represented in structural formulas by the letter R. A residue is a monomer that results when two or more amino acids combine and remove water molecules. The primary structure of a protein, a peptide chain, is made of amino acid residues. The unique characteristics of the functional groups and R groups allow these components of the amino acids to form hydrogen, ionic, and disulfide bonds, along with polar/nonpolar interactions needed to form secondary, tertiary, and quaternary protein structures. These groups are composed primarily of carbon, hydrogen, oxygen, nitrogen, and sulfur, in the form of hydrocarbons, acids, amides, alcohols, and amines. A few examples illustrating these possibilities are shown below.

A six-panel table titled 'Some Amino Acids and Their Structures.' Each panel draws an amino acid as a central carbon bonded to an amino group, a hydrogen, a carboxyl group, and a shaded side (R) group that differs by amino acid: lysine, glutamine, aspartate, serine, cysteine, and alanine.
Extended description

Reading left to right, top row then bottom row: lysine’s R group is a four-carbon chain ending in an amino group; glutamine’s is a two-carbon chain ending in a carbon double-bonded to oxygen and single-bonded to an amino group; aspartate’s is a one-carbon group ending in a second carboxyl group; serine’s is a one-carbon group ending in a hydroxyl group; cysteine’s is a one-carbon group ending in a sulfhydryl group; and alanine’s R group is a single methyl group. A note below the table states that blue shading marks the R group in each structure.

Amino acids may chemically bond together by reaction of the carboxylic acid group of one molecule with the amine group of another. This reaction forms a peptide bond and a water molecule and is another example of dehydration synthesis, shown below. Molecules formed by chemically linking relatively modest numbers of amino acids (approximately 50 or fewer) are called peptides, and prefixes are often used to specify these numbers: dipeptides (two amino acids), tripeptides (three amino acids), and so forth. More generally, the approximate number of amino acids is designated: oligopeptides are formed by joining up to approximately 20 amino acids, whereas polypeptides are synthesized from up to approximately 50 amino acids. When the number of amino acids linked together becomes very large, or when multiple polypeptides are used as building subunits, the macromolecules that result are called proteins. The continuously variable length (the number of monomers) of these biopolymers, along with the variety of possible R groups on each amino acid, allows for a nearly unlimited diversity in the types of proteins that may be formed.

Two separate alanine molecules, each drawn as H2N–CH(CH3)–COOH, react to release a water molecule (labeled H2O) and form a dipeptide: the two alanine units joined by a carbon-nitrogen bond, labeled 'peptide bond,' where the carboxyl carbon of the first alanine bonds to the amino nitrogen of the second.
Peptide bond formation is a dehydration synthesis reaction. The carboxyl group of the first amino acid (alanine) is linked to the amino group of the incoming second amino acid (alanine). In the process, a molecule of water is released.

Check Your Understanding

How many amino acids are in polypeptides?

Protein Structure

The size (length) and specific amino acid sequence of a protein are major determinants of its shape, and the shape of a protein is critical to its function. For example, in the process of biological nitrogen fixation (see the chapter Biogeochemical Cycles), soil microorganisms collectively known as rhizobia symbiotically interact with roots of legume plants such as soybeans, peanuts, or beans to form a novel structure called a nodule on the plant roots. The plant then produces a carrier protein called leghemoglobin, a protein that carries nitrogen or oxygen. Leghemoglobin binds with a very high affinity to its substrate oxygen at a specific region of the protein where the shape and amino acid sequence are appropriate (the active site). If the shape or chemical environment of the active site is altered, even slightly, the substrate may not be able to bind as strongly, or it may not bind at all. Thus, for the protein to be fully active, it must have the appropriate shape for its function.

Protein structure is categorized in terms of four levels: primary, secondary, tertiary, and quaternary. The primary structure is simply the sequence of amino acids that make up the polypeptide chain. The figure below depicts the primary structure of a protein.

A winding chain of colored circles representing linked amino acids, running from a free amino group (N-terminus) at one end to a free carboxyl group (C-terminus) at the other, with an inset showing a single amino acid's generic structure.
The primary structure of a protein is the sequence of amino acids. (credit: modification of work by National Human Genome Research Institute)
Extended description

Reading left to right: the chain begins at the free amino group, the N-terminus. Two adjoining circles partway along are labeled ‘amino acids,’ and the bond joining them is labeled ‘peptide bonds.’ The chain ends at the free carboxyl group, the C-terminus, where an inset zooms in on one residue’s structure — a central carbon bonded to an amino group, a hydrogen, an R group, and an acidic carboxyl group — and names the chain’s last four residues, in order, as Phe, Leu, Set, and Cys (the artwork prints Set where the serine abbreviation Ser is evidently meant).

The chain of amino acids that defines a protein’s primary structure is not rigid, but instead is flexible because of the nature of the bonds that hold the amino acids together. When the chain is sufficiently long, hydrogen bonding may occur between amine and carbonyl functional groups within the peptide backbone (excluding the R side group), resulting in localized folding of the polypeptide chain into helices and sheets. These shapes constitute a protein’s secondary structure. The most common secondary structures are the α-helix and β-pleated sheet. In the α-helix structure, the helix is held by hydrogen bonds between the oxygen atom in a carbonyl group of one amino acid and the hydrogen atom of the amino group that is just four amino acid units farther along the chain. In the β-pleated sheet, the pleats are formed by similar hydrogen bonds between continuous sequences of carbonyl and amino groups that are further separated on the backbone of the polypeptide chain, shown below.

A chain of spheres forms a coiled spiral labeled α-helix, and the same kind of chain also forms a ribbon that folds back on itself, labeled β-pleated sheet; two close-up insets show dotted lines marking hydrogen bonds between amino acids that hold each shape together.
The secondary structure of a protein may be an α-helix or a β-pleated sheet, or both.

The next level of protein organization is the tertiary structure, which is the large-scale three-dimensional shape of a single polypeptide chain. Tertiary structure is determined by interactions between amino acid residues that are far apart in the chain. A variety of interactions give rise to protein tertiary structure, such as disulfide bridges, which are bonds between the sulfhydryl (–SH) functional groups on amino acid side groups; hydrogen bonds; ionic bonds; and hydrophobic interactions between nonpolar side chains. All these interactions, weak and strong, combine to determine the final three-dimensional shape of the protein and its function, shown below.

A looping red ribbon labeled 'polypeptide backbone' with four close-ups of the interactions that fold it, three of them boxed: an ionic bond between a positively charged and a negatively charged side chain, hydrophobic interactions among nonpolar side chains, a disulfide linkage between two sulfur atoms, and a hydrogen bond between two polar side chains.
The tertiary structure of proteins is determined by a variety of attractive forces, including hydrophobic interactions, ionic bonding, hydrogen bonding, and disulfide linkages.
Extended description

Reading around the backbone: an ionic bond forms where a side chain ending in a positive charge pairs with one ending in a negative charge; a shaded box shows hydrophobic interactions between two branched, all-carbon-and-hydrogen side chains; a disulfide linkage is drawn as a sulfur atom on one loop bonded to a sulfur atom on a neighboring loop; and a hydrogen bond is drawn as a dotted line between two nearby polar side chains.

The process by which a polypeptide chain assumes a large-scale, three-dimensional shape is called protein folding. Folded proteins that are fully functional in their normal biological role are said to possess a native structure. When a protein loses its three-dimensional shape, it may no longer be functional. These unfolded proteins are denatured. Denaturation implies the loss of the secondary structure and tertiary structure (and, if present, the quaternary structure) without the loss of the primary structure.

Some proteins are assemblies of several separate polypeptides, also known as protein subunits. These proteins function adequately only when all subunits are present and appropriately configured. The interactions that hold these subunits together constitute the quaternary structure of the protein. The overall quaternary structure is stabilized by relatively weak interactions. Hemoglobin, for example, has a quaternary structure of four globular protein subunits: two α and two β polypeptides, each one containing an iron-based heme, shown below.

A four-lobed ribbon structure of coiled and wound ribbons in two colors, labeled α1, α2 (gold) and β1, β2 (green); a cluster of reddish spheres embedded in each lobe is labeled heme group.
A hemoglobin molecule has two α and two β polypeptides together with four heme groups.

Another important class of proteins is the conjugated proteins that have a nonprotein portion. If the conjugated protein has a carbohydrate attached, it is called a glycoprotein. If it has a lipid attached, it is called a lipoprotein. These proteins are important components of membranes. The figure below summarizes the four levels of protein structure.

Four labeled panels, left to right: 'Primary Protein Structure' shows a beaded chain of amino acids; 'Secondary Protein Structure' shows a spiral (α-helix) and a folded ribbon (β-pleated sheet); 'Tertiary Protein Structure' shows helices and sheets folded into one compact 3-D shape; 'Quaternary Protein Structure' shows two such folded shapes joined together.
Protein structure has four levels of organization. (credit: modification of work by National Human Genome Research Institute)

Check Your Understanding

What can happen if a protein’s primary, secondary, tertiary, or quaternary structure is changed?

Micro Connection. Primary Structure, Dysfunctional Proteins, and Cystic Fibrosis

Proteins associated with biological membranes are classified as extrinsic or intrinsic. Extrinsic proteins, also called peripheral proteins, are loosely associated with one side of the membrane. Intrinsic proteins, or integral proteins, are embedded in the membrane and often function as part of transport systems as transmembrane proteins. Cystic fibrosis (CF) is a human genetic disorder caused by a change in the transmembrane protein. It affects mostly the lungs but may also affect the pancreas, liver, kidneys, and intestine. CF is caused by a loss of the amino acid phenylalanine in a cystic fibrosis transmembrane protein (CFTR). The loss of one amino acid changes the primary structure of a protein that normally helps transport salt and water in and out of cells, shown below.

A phospholipid bilayer with two identical channel proteins spanning it. Chloride ions pass freely up through the channel on the right and out of the cell. A patch of mucus sits atop the channel on the left, and the chloride ions beneath it cannot pass through to the outside.
The normal CFTR protein is a channel protein that helps salt (sodium chloride) move in and out of cells.

The change in the primary structure prevents the protein from functioning properly, which causes the body to produce unusually thick mucus that clogs the lungs and leads to the accumulation of sticky mucus. The mucus obstructs the pancreas and stops natural enzymes from helping the body break down food and absorb vital nutrients.

In the lungs of individuals with cystic fibrosis, the altered mucus provides an environment where bacteria can thrive. This colonization leads to the formation of biofilms in the small airways of the lungs. The most common pathogens found in the lungs of patients with cystic fibrosis are Pseudomonas aeruginosa and Burkholderia cepacia, shown below. Pseudomonas differentiates within the biofilm in the lung and forms large colonies, called “mucoid” Pseudomonas. The colonies have a unique pigmentation that shows up in laboratory tests and provides physicians with the first clue that the patient has CF (such colonies are rare in healthy individuals).

(a) A colorized scanning electron micrograph of several tan, rod-shaped bacterial cells scattered among smaller pink spherical particles on a teal background. (b) A circular agar plate whose left side is covered in green-pigmented bacterial growth, with the green pigment visibly diffusing into the clear agar beyond the edge of the growth.
(a) A scanning electron micrograph shows the opportunistic bacterium Pseudomonas aeruginosa. (b) Pigment-producing P. aeruginosa on cetrimide agar shows the green pigment called pyocyanin. (credit a: modification of work by the Centers for Disease Control and Prevention)

Link to Learning

For more information about cystic fibrosis, visit the Cystic Fibrosis Foundation website.

Summary

  • Amino acids are small molecules essential to all life. Each has an α carbon to which a hydrogen atom, carboxyl group, and amine group are bonded. The fourth bonded group, represented by R, varies in chemical composition, size, polarity, and charge among different amino acids, providing variation in properties.
  • Peptides are polymers formed by the linkage of amino acids via dehydration synthesis. The bonds between the linked amino acids are called peptide bonds. The number of amino acids linked together may vary from a few to many.
  • Proteins are polymers formed by the linkage of a very large number of amino acids. They perform many important functions in a cell, serving as nutrients and enzymes; storage molecules for carbon, nitrogen, and energy; and structural components.
  • The structure of a protein is a critical determinant of its function and is described by a graduated classification: primary, secondary, tertiary, and quaternary. The native structure of a protein may be disrupted by denaturation, resulting in loss of its higher-order structure and its biological function.
  • Some proteins are formed by several separate protein subunits, the interaction of these subunits composing the quaternary structure of the protein complex.
  • Conjugated proteins have a nonpolypeptide portion that can be a carbohydrate (forming a glycoprotein) or a lipid fraction (forming a lipoprotein). These proteins are important components of membranes.

Key terms

  • amino acid — a molecule consisting of a hydrogen atom, a carboxyl group, and an amine group bonded to the same carbon. The group bonded to the carbon varies and is represented by an R in the structural formula.
  • side chain — the variable functional group, R, attached to the α carbon of an amino acid.
  • peptide bond — bond between the carboxyl group of one amino acid and the amine group of another; formed with the loss of a water molecule.
  • oligopeptides — peptide having up to approximately 20 amino acids.
  • polypeptides — polymer having from approximately 20 to 50 amino acids.
  • proteins — macromolecule that results when the number of amino acids linked together becomes very large, or when multiple polypeptides are used as building subunits.
  • primary structure — bonding sequence of amino acids in a polypeptide chain.
  • secondary structure — structure stabilized by hydrogen bonds between the carbonyl and amine groups of a polypeptide chain; may be an α-helix or a β-pleated sheet, or both.
  • α-helix — secondary structure consisting of a helix stabilized by hydrogen bonds between nearby amino acid residues in a polypeptide.
  • β-pleated sheet — secondary structure consisting of pleats formed by hydrogen bonds between localized segments of amino acid residues on the backbone of the polypeptide chain.
  • tertiary structure — large-scale, three-dimensional structure of a polypeptide.
  • disulfide bridges — covalent bond between the sulfur atoms of two sulfhydryl side chains.
  • native structure — three-dimensional structure of folded fully functional proteins.
  • denatured — describing a protein that has lost its secondary and tertiary structure (and quaternary structure, if applicable) without the loss of its primary structure.
  • quaternary structure — structure of protein complexes formed by the combination of several separate polypeptides or subunits.
  • conjugated proteins — protein carrying a nonpolypeptidic portion.
  • glycoprotein — conjugated protein with a carbohydrate attached.
  • lipoprotein — conjugated protein attached to a lipid.

Practice

Describe the fundamental structure of an amino acid

Which of the following groups varies among different amino acids?

The amino acids present in proteins differ in which of the following?

A four-unit peptide chain drawn as a structural formula, colored green at the left end and blue at the right end with two black units in between; each unit's carbon-oxygen group connects to the nitrogen-hydrogen group of the next unit, chaining the four units together.
A structural formula of a tetrapeptide made of four amino acid units joined end to end.

The image above represents a tetrapeptide. How many peptide bonds are in this molecule?

Identify the side groups of the four amino acids composing this tetrapeptide.

Show model answer
Reading the structure from left (green) to right (blue): the first amino acid’s side chain is a single carbon bonded to two methyl groups; the second amino acid’s side chain has no additional atoms beyond a hydrogen; the third amino acid’s side chain is a one-carbon group ending in a hydroxyl group; and the fourth amino acid’s side chain is a single methyl group.

Did your answer mention:

Describe the chemical structures of proteins

Which of the following bonds are not involved in tertiary structure?

The sequence of amino acids in a protein is called its ________.

Denaturation implies the loss of the ________ and ________ structures without the loss of the ________ structure.

Summarize the unique characteristics of proteins

A change in one amino acid in a protein sequence always results in a loss of function.

Heating a protein sufficiently may cause it to denature. Considering the definition of denaturation, what does this statement say about the strengths of peptide bonds in comparison to hydrogen bonds?

Show model answer
Heating a protein is enough to break the hydrogen bonds and other weak interactions that hold its secondary and tertiary structure together, denaturing it — but denaturation does not break the peptide bonds of the primary structure, since the primary structure survives intact. Because the same amount of heat breaks the hydrogen bonds but not the peptide bonds, peptide bonds must be stronger than hydrogen bonds.

Did your answer mention:

A protein that carries a nonprotein portion — a carbohydrate or a lipid — is called a ________ protein.

This section is adapted from Microbiology, Section 7.4: Proteins by Nina Parker, Mark Schneegurt, Anh-Hue Thi Tu, Philip Lister, Brian M. Forster, and OpenStax, © OpenStax, licensed under CC BY-NC-SA 4.0. Access the original for free at openstax.org. Changes: all 10 figures re-encoded as WebP, with kind set explicitly after inspection — diagram for the nine structural, cartoon, and rendered-model figures (the amino-acid table, the peptide-bond reaction, and the primary, secondary, tertiary, hemoglobin, four-level-summary, CFTR-channel, and tetrapeptide diagrams) despite the media manifest’s JPEG-based guess of “photo” for all of them, and photo for the two-panel micrograph-and-plate figure, whose panels are genuine photographs; alts rewritten from the served images rather than the source’s alt text, and a longdesc added to the amino-acid table, the primary-structure diagram, and the tertiary-structure diagram, each walking its labeled parts in reading order (the primary-structure artwork prints a residue label “Set” where “Ser” is meant; the longdesc reports the printed label and notes it); the source alt for the amino-acid table is corrected in the rewrite — it names only five of the image’s six amino acids (omitting glutamine) and describes an ionized (zwitterionic) form the image does not draw (both reported as source-alt defects); the five Protein Structure figures, which the source floats together as one block after their five paragraphs for print pagination, are each placed at the paragraph that first introduces it, for a scrolling page; the Micro Connection box’s second, repeated cross-reference to the Pseudomonas micrograph-and-plate figure is merged into its first mention, since the same figure cannot appear twice on the page; the Micro Connection and Link to Learning notes are rendered as callouts, the Micro Connection keeping its source title and the Link to Learning keeping its URL and sentence boundary; the cross-reference to Biogeochemical Cycles (m58825, chapter 8) is left as plain text because that chapter is not yet authored; both of the module’s body Check Your Understanding questions are graded here as multiplechoice items, each fixed by one sentence of this module (the oligopeptide/polypeptide size-range sentence, and the sentence stating that an unfolded protein becomes denatured and may lose function); of the module’s two unkeyed Critical Thinking questions (this module prints no Short Answer set), the tetrapeptide question splits into its two parts as the source prints them: part (a), the peptide-bond count, is graded as a multiplechoice keyed from the figure (the honest caption and alt describe the chain’s connectivity without stating the count), and part (b), the side groups, stays a self-check whose model answer is likewise read from the figure; the other Critical Thinking question, comparing the strength of peptide bonds to hydrogen bonds via denaturation, stays a self-check because the comparison is an inference from this module’s denaturation and tertiary-structure sentences rather than a fact either sentence states outright; the three Multiple Choice, one Fill in the Blank, and one True/False items keep the source’s own keys, options, and order; the module’s other Fill in the Blank item has three blanks whose filled order is the point (“secondary, tertiary, primary”), so — following this book’s rule for an ordered multi-blank item — it is rebuilt as a multiplechoice among that triplet and three other orderings of the module’s own four structural-level terms, rather than a textin, because a single text field cannot grade three separate blanks; one term-recall textin (“conjugated protein”) is added from Key terms to round out the third objective group, which the source’s own exercise set leaves at two items; no source exercise was omitted — all eight of the module’s exercises are used; key terms compiled from the module’s 18 distinct defined terms and the book’s Glossary appendix; none needed a sentence-derived definition. The appendix’s own “primary structure” entry is a headword-loss defect of the same class already on record (erratum 320): its single <item> merges a second, headword-less entry for “protein” as a run-on tail (“bonding sequence of amino acids in a polypeptide chain protein macromolecule that results when…”); this page’s primary structure bullet uses only the first clause, and the merged tail is recovered verbatim as this page’s proteins bullet, rather than the sentence-derived definition a naive scaffold would produce (reported as a new source defect); the denatured bullet uses the nearest-headword entry, “denatured protein,” whose definition matches this module’s sense of an unfolded, nonfunctional protein.