Proteins
By the end of this section, you will be able to:
- Describe the fundamental structure of an amino acid
- Describe the chemical structures of proteins
- Summarize the unique characteristics of proteins
At the beginning of this chapter, a famous experiment was described in which scientists synthesized amino acids under conditions simulating those present on earth long before the evolution of life as we know it. These compounds are capable of bonding together in essentially any number, yielding molecules of essentially any size that possess a wide array of physical and chemical properties and perform numerous functions vital to all organisms. The molecules derived from amino acids can function as structural components of cells and subcellular entities, as sources of nutrients, as atom- and energy-storage reservoirs, and as functional species such as hormones, enzymes, receptors, and transport molecules.
Amino Acids and Peptide Bonds
An amino acid is an organic molecule in which a hydrogen atom, a carboxyl group (–COOH), and an amino group (–NH₂) are all bonded to the same carbon atom, the so-called α carbon. The fourth group bonded to the α carbon varies among the different amino acids and is called a residue or a side chain, represented in structural formulas by the letter R. A residue is a monomer that results when two or more amino acids combine and remove water molecules. The primary structure of a protein, a peptide chain, is made of amino acid residues. The unique characteristics of the functional groups and R groups allow these components of the amino acids to form hydrogen, ionic, and disulfide bonds, along with polar/nonpolar interactions needed to form secondary, tertiary, and quaternary protein structures. These groups are composed primarily of carbon, hydrogen, oxygen, nitrogen, and sulfur, in the form of hydrocarbons, acids, amides, alcohols, and amines. A few examples illustrating these possibilities are shown below.

Extended description
Reading left to right, top row then bottom row: lysine’s R group is a four-carbon chain ending in an amino group; glutamine’s is a two-carbon chain ending in a carbon double-bonded to oxygen and single-bonded to an amino group; aspartate’s is a one-carbon group ending in a second carboxyl group; serine’s is a one-carbon group ending in a hydroxyl group; cysteine’s is a one-carbon group ending in a sulfhydryl group; and alanine’s R group is a single methyl group. A note below the table states that blue shading marks the R group in each structure.
Amino acids may chemically bond together by reaction of the carboxylic acid group of one molecule with the amine group of another. This reaction forms a peptide bond and a water molecule and is another example of dehydration synthesis, shown below. Molecules formed by chemically linking relatively modest numbers of amino acids (approximately 50 or fewer) are called peptides, and prefixes are often used to specify these numbers: dipeptides (two amino acids), tripeptides (three amino acids), and so forth. More generally, the approximate number of amino acids is designated: oligopeptides are formed by joining up to approximately 20 amino acids, whereas polypeptides are synthesized from up to approximately 50 amino acids. When the number of amino acids linked together becomes very large, or when multiple polypeptides are used as building subunits, the macromolecules that result are called proteins. The continuously variable length (the number of monomers) of these biopolymers, along with the variety of possible R groups on each amino acid, allows for a nearly unlimited diversity in the types of proteins that may be formed.

Check Your Understanding
How many amino acids are in polypeptides?
Compare the size ranges the module gives for oligopeptides, polypeptides, and proteins.Protein Structure
The size (length) and specific amino acid sequence of a protein are major determinants of its shape, and the shape of a protein is critical to its function. For example, in the process of biological nitrogen fixation (see the chapter Biogeochemical Cycles), soil microorganisms collectively known as rhizobia symbiotically interact with roots of legume plants such as soybeans, peanuts, or beans to form a novel structure called a nodule on the plant roots. The plant then produces a carrier protein called leghemoglobin, a protein that carries nitrogen or oxygen. Leghemoglobin binds with a very high affinity to its substrate oxygen at a specific region of the protein where the shape and amino acid sequence are appropriate (the active site). If the shape or chemical environment of the active site is altered, even slightly, the substrate may not be able to bind as strongly, or it may not bind at all. Thus, for the protein to be fully active, it must have the appropriate shape for its function.
Protein structure is categorized in terms of four levels: primary, secondary, tertiary, and quaternary. The primary structure is simply the sequence of amino acids that make up the polypeptide chain. The figure below depicts the primary structure of a protein.

Extended description
Reading left to right: the chain begins at the free amino group, the N-terminus. Two adjoining circles partway along are labeled ‘amino acids,’ and the bond joining them is labeled ‘peptide bonds.’ The chain ends at the free carboxyl group, the C-terminus, where an inset zooms in on one residue’s structure — a central carbon bonded to an amino group, a hydrogen, an R group, and an acidic carboxyl group — and names the chain’s last four residues, in order, as Phe, Leu, Set, and Cys (the artwork prints Set where the serine abbreviation Ser is evidently meant).
The chain of amino acids that defines a protein’s primary structure is not rigid, but instead is flexible because of the nature of the bonds that hold the amino acids together. When the chain is sufficiently long, hydrogen bonding may occur between amine and carbonyl functional groups within the peptide backbone (excluding the R side group), resulting in localized folding of the polypeptide chain into helices and sheets. These shapes constitute a protein’s secondary structure. The most common secondary structures are the α-helix and β-pleated sheet. In the α-helix structure, the helix is held by hydrogen bonds between the oxygen atom in a carbonyl group of one amino acid and the hydrogen atom of the amino group that is just four amino acid units farther along the chain. In the β-pleated sheet, the pleats are formed by similar hydrogen bonds between continuous sequences of carbonyl and amino groups that are further separated on the backbone of the polypeptide chain, shown below.

The next level of protein organization is the tertiary structure, which is the large-scale three-dimensional shape of a single polypeptide chain. Tertiary structure is determined by interactions between amino acid residues that are far apart in the chain. A variety of interactions give rise to protein tertiary structure, such as disulfide bridges, which are bonds between the sulfhydryl (–SH) functional groups on amino acid side groups; hydrogen bonds; ionic bonds; and hydrophobic interactions between nonpolar side chains. All these interactions, weak and strong, combine to determine the final three-dimensional shape of the protein and its function, shown below.

Extended description
Reading around the backbone: an ionic bond forms where a side chain ending in a positive charge pairs with one ending in a negative charge; a shaded box shows hydrophobic interactions between two branched, all-carbon-and-hydrogen side chains; a disulfide linkage is drawn as a sulfur atom on one loop bonded to a sulfur atom on a neighboring loop; and a hydrogen bond is drawn as a dotted line between two nearby polar side chains.
The process by which a polypeptide chain assumes a large-scale, three-dimensional shape is called protein folding. Folded proteins that are fully functional in their normal biological role are said to possess a native structure. When a protein loses its three-dimensional shape, it may no longer be functional. These unfolded proteins are denatured. Denaturation implies the loss of the secondary structure and tertiary structure (and, if present, the quaternary structure) without the loss of the primary structure.
Some proteins are assemblies of several separate polypeptides, also known as protein subunits. These proteins function adequately only when all subunits are present and appropriately configured. The interactions that hold these subunits together constitute the quaternary structure of the protein. The overall quaternary structure is stabilized by relatively weak interactions. Hemoglobin, for example, has a quaternary structure of four globular protein subunits: two α and two β polypeptides, each one containing an iron-based heme, shown below.

Another important class of proteins is the conjugated proteins that have a nonprotein portion. If the conjugated protein has a carbohydrate attached, it is called a glycoprotein. If it has a lipid attached, it is called a lipoprotein. These proteins are important components of membranes. The figure below summarizes the four levels of protein structure.

Check Your Understanding
What can happen if a protein’s primary, secondary, tertiary, or quaternary structure is changed?
The passage right after quaternary structure explains what happens when a protein unfolds.Micro Connection. Primary Structure, Dysfunctional Proteins, and Cystic Fibrosis
Proteins associated with biological membranes are classified as extrinsic or intrinsic. Extrinsic proteins, also called peripheral proteins, are loosely associated with one side of the membrane. Intrinsic proteins, or integral proteins, are embedded in the membrane and often function as part of transport systems as transmembrane proteins. Cystic fibrosis (CF) is a human genetic disorder caused by a change in the transmembrane protein. It affects mostly the lungs but may also affect the pancreas, liver, kidneys, and intestine. CF is caused by a loss of the amino acid phenylalanine in a cystic fibrosis transmembrane protein (CFTR). The loss of one amino acid changes the primary structure of a protein that normally helps transport salt and water in and out of cells, shown below.

The change in the primary structure prevents the protein from functioning properly, which causes the body to produce unusually thick mucus that clogs the lungs and leads to the accumulation of sticky mucus. The mucus obstructs the pancreas and stops natural enzymes from helping the body break down food and absorb vital nutrients.
In the lungs of individuals with cystic fibrosis, the altered mucus provides an environment where bacteria can thrive. This colonization leads to the formation of biofilms in the small airways of the lungs. The most common pathogens found in the lungs of patients with cystic fibrosis are Pseudomonas aeruginosa and Burkholderia cepacia, shown below. Pseudomonas differentiates within the biofilm in the lung and forms large colonies, called “mucoid” Pseudomonas. The colonies have a unique pigmentation that shows up in laboratory tests and provides physicians with the first clue that the patient has CF (such colonies are rare in healthy individuals).

Link to Learning
For more information about cystic fibrosis, visit the Cystic Fibrosis Foundation website.
Summary
- Amino acids are small molecules essential to all life. Each has an α carbon to which a hydrogen atom, carboxyl group, and amine group are bonded. The fourth bonded group, represented by R, varies in chemical composition, size, polarity, and charge among different amino acids, providing variation in properties.
- Peptides are polymers formed by the linkage of amino acids via dehydration synthesis. The bonds between the linked amino acids are called peptide bonds. The number of amino acids linked together may vary from a few to many.
- Proteins are polymers formed by the linkage of a very large number of amino acids. They perform many important functions in a cell, serving as nutrients and enzymes; storage molecules for carbon, nitrogen, and energy; and structural components.
- The structure of a protein is a critical determinant of its function and is described by a graduated classification: primary, secondary, tertiary, and quaternary. The native structure of a protein may be disrupted by denaturation, resulting in loss of its higher-order structure and its biological function.
- Some proteins are formed by several separate protein subunits, the interaction of these subunits composing the quaternary structure of the protein complex.
- Conjugated proteins have a nonpolypeptide portion that can be a carbohydrate (forming a glycoprotein) or a lipid fraction (forming a lipoprotein). These proteins are important components of membranes.
Key terms
- amino acid — a molecule consisting of a hydrogen atom, a carboxyl group, and an amine group bonded to the same carbon. The group bonded to the carbon varies and is represented by an R in the structural formula.
- side chain — the variable functional group, R, attached to the α carbon of an amino acid.
- peptide bond — bond between the carboxyl group of one amino acid and the amine group of another; formed with the loss of a water molecule.
- oligopeptides — peptide having up to approximately 20 amino acids.
- polypeptides — polymer having from approximately 20 to 50 amino acids.
- proteins — macromolecule that results when the number of amino acids linked together becomes very large, or when multiple polypeptides are used as building subunits.
- primary structure — bonding sequence of amino acids in a polypeptide chain.
- secondary structure — structure stabilized by hydrogen bonds between the carbonyl and amine groups of a polypeptide chain; may be an α-helix or a β-pleated sheet, or both.
- α-helix — secondary structure consisting of a helix stabilized by hydrogen bonds between nearby amino acid residues in a polypeptide.
- β-pleated sheet — secondary structure consisting of pleats formed by hydrogen bonds between localized segments of amino acid residues on the backbone of the polypeptide chain.
- tertiary structure — large-scale, three-dimensional structure of a polypeptide.
- disulfide bridges — covalent bond between the sulfur atoms of two sulfhydryl side chains.
- native structure — three-dimensional structure of folded fully functional proteins.
- denatured — describing a protein that has lost its secondary and tertiary structure (and quaternary structure, if applicable) without the loss of its primary structure.
- quaternary structure — structure of protein complexes formed by the combination of several separate polypeptides or subunits.
- conjugated proteins — protein carrying a nonpolypeptidic portion.
- glycoprotein — conjugated protein with a carbohydrate attached.
- lipoprotein — conjugated protein attached to a lipid.
Practice
Describe the fundamental structure of an amino acid
Which of the following groups varies among different amino acids?
Compare which parts of the structure are the same in every amino acid and which one differs.The amino acids present in proteins differ in which of the following?
Check whether more than one of the listed properties varies among amino acids.
The image above represents a tetrapeptide. How many peptide bonds are in this molecule?
Count the amide linkages that join the four amino acid units end to end.Identify the side groups of the four amino acids composing this tetrapeptide.
Show model answer
Did your answer mention:
Describe the chemical structures of proteins
Which of the following bonds are not involved in tertiary structure?
Tertiary structure is held together by interactions between side chains, not by the bonds within the backbone.The sequence of amino acids in a protein is called its ________.
This is the first of the four levels of protein organization described in this section.Denaturation implies the loss of the ________ and ________ structures without the loss of the ________ structure.
Re-read the denaturation sentence for which structural level survives unfolding intact.Summarize the unique characteristics of proteins
A change in one amino acid in a protein sequence always results in a loss of function.
Compare this claim with how the module describes variation among amino acid side groups and their role in protein activity.Heating a protein sufficiently may cause it to denature. Considering the definition of denaturation, what does this statement say about the strengths of peptide bonds in comparison to hydrogen bonds?
Show model answer
Did your answer mention:
A protein that carries a nonprotein portion — a carbohydrate or a lipid — is called a ________ protein.
This section’s Key terms list the term for a protein with an attached nonprotein portion.This section is adapted from Microbiology, Section 7.4: Proteins by Nina Parker, Mark Schneegurt, Anh-Hue Thi Tu, Philip Lister, Brian M. Forster, and OpenStax, © OpenStax, licensed under CC BY-NC-SA 4.0. Access the original for free at openstax.org. Changes: all 10 figures re-encoded as WebP, with kind set explicitly after inspection — diagram for the nine structural, cartoon, and rendered-model figures (the amino-acid table, the peptide-bond reaction, and the primary, secondary, tertiary, hemoglobin, four-level-summary, CFTR-channel, and tetrapeptide diagrams) despite the media manifest’s JPEG-based guess of “photo” for all of them, and photo for the two-panel micrograph-and-plate figure, whose panels are genuine photographs; alts rewritten from the served images rather than the source’s alt text, and a longdesc added to the amino-acid table, the primary-structure diagram, and the tertiary-structure diagram, each walking its labeled parts in reading order (the primary-structure artwork prints a residue label “Set” where “Ser” is meant; the longdesc reports the printed label and notes it); the source alt for the amino-acid table is corrected in the rewrite — it names only five of the image’s six amino acids (omitting glutamine) and describes an ionized (zwitterionic) form the image does not draw (both reported as source-alt defects); the five Protein Structure figures, which the source floats together as one block after their five paragraphs for print pagination, are each placed at the paragraph that first introduces it, for a scrolling page; the Micro Connection box’s second, repeated cross-reference to the Pseudomonas micrograph-and-plate figure is merged into its first mention, since the same figure cannot appear twice on the page; the Micro Connection and Link to Learning notes are rendered as callouts, the Micro Connection keeping its source title and the Link to Learning keeping its URL and sentence boundary; the cross-reference to Biogeochemical Cycles (m58825, chapter 8) is left as plain text because that chapter is not yet authored; both of the module’s body Check Your Understanding questions are graded here as multiplechoice items, each fixed by one sentence of this module (the oligopeptide/polypeptide size-range sentence, and the sentence stating that an unfolded protein becomes denatured and may lose function); of the module’s two unkeyed Critical Thinking questions (this module prints no Short Answer set), the tetrapeptide question splits into its two parts as the source prints them: part (a), the peptide-bond count, is graded as a multiplechoice keyed from the figure (the honest caption and alt describe the chain’s connectivity without stating the count), and part (b), the side groups, stays a self-check whose model answer is likewise read from the figure; the other Critical Thinking question, comparing the strength of peptide bonds to hydrogen bonds via denaturation, stays a self-check because the comparison is an inference from this module’s denaturation and tertiary-structure sentences rather than a fact either sentence states outright; the three Multiple Choice, one Fill in the Blank, and one True/False items keep the source’s own keys, options, and order; the module’s other Fill in the Blank item has three blanks whose filled order is the point (“secondary, tertiary, primary”), so — following this book’s rule for an ordered multi-blank item — it is rebuilt as a multiplechoice among that triplet and three other orderings of the module’s own four structural-level terms, rather than a textin, because a single text field cannot grade three separate blanks; one term-recall textin (“conjugated protein”) is added from Key terms to round out the third objective group, which the source’s own exercise set leaves at two items; no source exercise was omitted — all eight of the module’s exercises are used; key terms compiled from the module’s 18 distinct defined terms and the book’s Glossary appendix; none needed a sentence-derived definition. The appendix’s own “primary structure” entry is a headword-loss defect of the same class already on record (erratum 320): its single <item> merges a second, headword-less entry for “protein” as a run-on tail (“bonding sequence of amino acids in a polypeptide chain protein macromolecule that results when…”); this page’s primary structure bullet uses only the first clause, and the merged tail is recovered verbatim as this page’s proteins bullet, rather than the sentence-derived definition a naive scaffold would produce (reported as a new source defect); the denatured bullet uses the nearest-headword entry, “denatured protein,” whose definition matches this module’s sense of an unfolded, nonfunctional protein.