Skip to content

Genomics and Proteomics

By the end of this section, you will be able to:

  • Explain systems biology
  • Describe a proteome
  • Define protein signature

Proteins are the final products of genes, which help perform the function that the gene encodes. Amino acids comprise proteins and play important roles in the cell. All enzymes (except ribozymes) are proteins that act as catalysts to affect the rate of reactions. Proteins are also regulatory molecules, and some are hormones. Transport proteins, such as hemoglobin, help transport oxygen to various organs. Antibodies that defend against foreign particles are also proteins. In the diseased state, protein function can be impaired because of changes at the genetic level or because of direct impact on a specific protein.

A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins. Proteomics is the study of proteomes’ function. Proteomics complements genomics and is useful when scientists want to test their hypotheses that they based on genes. Even though all multicellular organisms’ cells have the same set of genes, the set of proteins produced in different tissues is different and dependent on gene expression. Thus, the genome is constant, but the proteome varies and is dynamic within an organism. In addition, RNAs can be alternately spliced (cut and pasted to create novel combinations and novel proteins) and many proteins modify themselves after translation by processes such as proteolytic cleavage, phosphorylation, glycosylation, and ubiquitination. There are also protein-protein interactions, which complicate studying proteomes. Although the genome provides a blueprint, the final architecture depends on several factors that can change the progression of events that generate the proteome.

Metabolomics is related to genomics and proteomics. Metabolomics involves studying small molecule metabolites in an organism. The metabolome is the complete set of metabolites that are related to an organism’s genetic makeup. Metabolomics offers an opportunity to compare genetic makeup and physical characteristics, as well as genetic makeup and environmental factors. The goal of metabolome research is to identify, quantify, and catalogue all the metabolites in living organisms’ tissues and fluids.

Basic Techniques in Protein Analysis

The ultimate goal of proteomics is to identify or compare the proteins expressed from a given genome under specific conditions, study the interactions between the proteins, and use the information to predict cell behavior or develop drug targets. Just as scientists analyze the genome using the basic DNA sequencing technique, proteomics requires techniques for protein analysis. The basic technique for protein analysis, analogous to DNA sequencing, is mass spectrometry. Mass spectrometry identifies and determines a molecule’s characteristics. Advances in spectrometry have allowed researchers to analyze very small protein samples. X-ray crystallography, for example, enables scientists to determine a protein crystal’s three-dimensional structure at atomic resolution. Another protein imaging technique, nuclear magnetic resonance (NMR), uses atoms’ magnetic properties to determine the protein’s three-dimensional structure in aqueous solution. Scientists have also used protein microarrays to study protein interactions. Large-scale adaptations of the basic two-hybrid screen (see the diagram below) have provided the basis for protein microarrays. Scientists use computer software to analyze the vast amount of data for proteomic analysis.

Genomic- and proteomic-scale analyses are part of systems biology, which is the study of whole biological systems (genomes and proteomes) based on interactions within the system. The European Bioinformatics Institute and the Human Proteome Organization (HUPO) are developing and establishing effective tools to sort through the enormous pile of systems biology data. Because proteins are the direct products of genes and reflect activity at the genomic level, it is natural to use proteomes to compare the protein profiles of different cells to identify proteins and genes involved in disease processes. Most pharmaceutical drug trials target proteins. Researchers use information that they obtain from proteomics to identify novel drugs and to understand their mechanisms of action.

A two-panel diagram of two-hybrid screening. Top: a bait protein attached to the DNA-binding domain (BD) and a prey protein attached to the activator domain (AD) interact, so BD and AD come together on the DNA and an arrow shows transcription reaching the reporter gene. Bottom: the prey and bait do not interact, BD and AD stay apart, and no arrow reaches the reporter gene, so no transcription occurs.
Scientists use two-hybrid screening to determine whether two proteins interact. In this method, a transcription factor splits into a DNA-binding domain (BD) and an activator domain (AD). The binding domain is able to bind the promoter in the activator domain’s absence, but it does not turn on transcription. The bait protein attaches to the BD, and the prey protein attaches to the AD. Transcription occurs only if the prey “catches” the bait.
Extended description

Two stacked panels, each showing an orange DNA bar with a green oval labeled BD bound to it, captioned ‘Transcriptional activator binding domain’ with a leader line to the DNA. Top panel: a yellow pentagon labeled Bait sits on BD, and a salmon-pink shape labeled Prey touching Bait is attached to a blue-gray box labeled AD; a black arrow runs from the DNA bar to a blue-gray box labeled Reporter gene; text below reads ‘If the bait protein interacts with the prey protein, the promoter’s activator domain binds to the binding domain, and transcription occurs.’ Bottom panel: BD sits on the DNA with no Bait pentagon; a separate Prey shape with its attached AD box sits apart, not touching BD; the Reporter gene box appears with no arrow leading to it; text below reads ‘If the prey doesn’t catch the bait no transcription occurs.’

Scientists are challenged when implementing proteomic analysis because it is difficult to detect small protein quantities. Although mass spectrometry is good for detecting small protein amounts, variations in protein expression in diseased states can be difficult to discern. Proteins are naturally unstable molecules, which makes proteomic analysis much more difficult than genomic analysis.

Cancer Proteomics

Researchers are studying patients’ genomes and proteomes to understand the genetic basis of diseases. The most prominent disease researchers are studying with proteomic approaches is cancer. These approaches improve screening and early cancer detection. Researchers are able to identify proteins whose expression indicates the disease process. An individual protein is a biomarker; whereas, a set of proteins with altered expression levels is a protein signature. For a biomarker or protein signature to be useful as a candidate for early cancer screening and detection, they must secrete in body fluids, such as sweat, blood, or urine, such that health professionals can perform large-scale screenings in a noninvasive fashion. The current problem with using biomarkers for early cancer detection is the high rate of false-negative results. A false negative is an incorrect test result that should have been positive. In other words, many cancer cases go undetected, which makes biomarkers unreliable. Some examples of protein biomarkers in cancer detection are CA-125 for ovarian cancer and PSA for prostate cancer. Protein signatures may be more reliable than biomarkers to detect cancer cells. Researchers are also using proteomics to develop individualized treatment plans, which involves predicting whether or not an individual will respond to specific drugs and the side effects that the individual may experience. Researchers also use proteomics to predict the possibility of disease recurrence.

The National Cancer Institute has developed programs to improve cancer detection and treatment. The Clinical Proteomic Technologies for Cancer and the Early Detection Research Network are efforts to identify protein signatures specific to different cancer types. The Biomedical Proteomics Program identifies protein signatures and designs effective therapies for cancer patients.

Summary

Proteomics is the study of the entire set of proteins expressed by a given type of cell under certain environmental conditions. In a multicellular organism, different cell types will have different proteomes, and these will vary with environmental changes. Unlike a genome, a proteome is dynamic and in constant flux, which makes it both more complicated and more useful than the knowledge of genomes alone.

Proteomics approaches rely on protein analysis. Researchers are constantly upgrading these techniques. Researchers have used proteomics to study different cancer types. Medical professionals are using different biomarkers and protein signatures to analyze each cancer type. The future goal is to have a personalized treatment plan for each individual.

Key terms

  • biomarker — individual protein that is uniquely produced in a diseased state
  • false negative — incorrect test result that should have been positive
  • metabolome — complete set of metabolites which are related to an organism’s genetic makeup
  • metabolomics — study of small molecule metabolites in an organism
  • protein signature — set of uniquely expressed proteins in the diseased state
  • proteome — entire set of proteins that cell type produces
  • proteomics — study of proteomes’ function
  • systems biology — study of whole biological systems (genomes and proteomes) based on interactions within the system

Practice

Explain systems biology

The study of whole biological networks—genomes and proteomes together—based on how their parts interact is called ________.

Proteomics approaches rely on ________.

What is systems biology, and how do proteomics and genomics research relate to it?

Show model answer
Systems biology is the study of whole biological systems—genomes and proteomes—based on the interactions within the system. Genomic- and proteomic-scale analyses are part of systems biology, and researchers use that combined information to compare protein profiles between cells and identify proteins and genes involved in disease processes.

Did your answer mention:

Describe a proteome

The entire set of proteins that a cell type produces is called a(n) ________.

The study of a proteome’s function is called ________.

The study of small molecule metabolites in an organism is called ________.

The complete set of metabolites related to an organism’s genetic makeup is called the ________.

Define protein signature

What is a biomarker?

A protein signature is:

An individual protein that is uniquely produced in a diseased state is called a(n) ________.

A set of proteins with altered expression levels in a diseased state is called a(n) ________.

An incorrect test result that should have been positive is called a(n) ________.

How has proteomics been used in cancer detection and treatment?

Show model answer
Proteomics has provided a way to detect biomarkers and protein signatures, which have been used to screen for the early detection of cancer.

Did your answer mention:

What is personalized medicine?

Show model answer
Personalized medicine is the use of an individual’s genomic sequence to predict the risk for specific diseases. When a disease does occur, it can be used to develop a personalized treatment plan.

Did your answer mention:


This section is adapted from Biology 2e, Section 17.5: Genomics and Proteomics by Mary Ann Clark, Jung Choi, Matthew Douglas, and OpenStax, © OpenStax, licensed under CC BY-NC-SA 4.0. Access the original for free at openstax.org. Changes: the figure re-encoded as WebP and re-kinded from the manifest’s file-extension guess of “photo” to “diagram,” since it is a labeled schematic, not a photograph; its source alt’s letter-spaced “D N A binding domain” rewritten as “DNA-binding domain,” and a longdesc added since the two panels’ arrows, shapes, and in-image captions are not fully carried by the source caption; the body’s inline numbered cross-reference to the figure changed to a descriptive “see the diagram below,” since figures are not numbered here; a source typo (“proteoms” for “proteomes,” confirmed against the printed edition) silently corrected; the end-of-section Review Questions and Critical Thinking Questions adapted into the closing interactive Practice block (multiple choice and self-check respectively); eight key-term recall items added from the glossary; rubric checkpoints added to every self-check, decomposing its model answer (the source solution, or — for the first objective’s locally written item — the section’s own sentences) into check-off clauses with no new claims; a summary-derived cloze textin item (“protein analysis”) added under the first objective; and one additional self-check written locally, paraphrasing the section’s own paragraph defining systems biology, since the module’s two Review Questions and two Critical Thinking Questions all map to the third objective and left the first objective (“Explain systems biology”) without enough coverage to meet this book’s practice floor.