Phylogenetic Analysis of Mass Spectrometry or Gene Array Data for the Diagnosis of Physiological Conditions
Abstract
A universal data-mining platform capable of analyzing mass spectrometry (MS) serum proteomic profiles and/or gene array data to produce biologically meaningful classification; i.e., group together biologically related specimens into clades. This platform utilizes the principles of phylogenetics, such as parsimony, to reveal susceptibility to cancer development (or other physiological or pathophysiological conditions), diagnosis and typing of cancer, identifying stages of cancer, as well as post-treatment evaluation. To place specimens into their corresponding clade(s), the invention utilizes two algorithms: a new data-mining parsing algorithm, and a publicly available phylogenetic algorithm (MIX). By outgroup comparison (i.e., using a normal set as the standard reference), the parsing algorithm identifies under and/or overexpressed gene values or in the case of sera, (i) novel or (ii) vanished MS peaks, and peaks signifying (iii) up or (iv) down regulated proteins, and scores the variations as either derived (do not exit in the outgroup set) or ancestral (exist in the outgroup set); the derived is given a score of “1”, and the ancestral a score of “0”—these are called the polarized values. Furthermore, the shared derived characters that it identifies are potential biomarkers for cancers and other conditions and their subclasses.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method of quantifying progression of disease in an individual comprising:
a) obtaining a first blood serum specimen, wherein the first blood serum specimen is a diseased specimen comprising cells from at least one mammal afflicted with a disease; b) obtaining a second blood serum specimen, wherein the second blood serum specimen is an outgroup comprising cells from at least one mammal not afflicted with the disease. c) determining m/z values by mass spectrometry for proteins contained in the diseased specimen and the outgroup; d) extracting an absolute minimum and an absolute maximum for each m/z value determined for the outgroup, wherein the absolute minimum is separated into a first vector and the absolute maximum is separated into a second vector; e) polarizing the m/z values of the diseased specimen with the m/z values of the outgroup to produce a cladogram; and f) interpreting the cladogram.
2 . The method according to claim 1 , wherein the disease is cancer.
3 . The method according to claim 1 , wherein the diseased specimen comprises cells obtained from mammals afflicted with cancer selected from the group consisting of: breast, colon, lung, liver, ovarian, prostate, and pancreatic.
4 . The method according to claim 1 , wherein a first score is assigned to m/z values of the diseased specimen within a range between the absolute minimum and the absolute maximum, and a second score is assigned to m/z values of the diseased specimen falling outside the range between the absolute minimum and the absolute maximum.
5 . The method according to claim 1 further comprising using a computer program to determine:
a) whether each m/z value of the diseased specimen correlates to protein expression within a range between the first vector and the second vector;
b) whether each m/z value of the diseased specimen correlates to a vanished peak that otherwise exists in the range between the first vector and the second vector;
c) whether each m/z value of the diseased specimen correlates to protein expression above the second vector; and
d) whether each m/z value of the diseased specimen correlates to protein expression below the first vector.
6 . A method of determining the progressive state of disease in an individual comprising:
a) obtaining a first blood serum specimen, wherein the first blood serum specimen is a diseased specimen comprising cells from at least one mammal afflicted with a disease; b) obtaining a second blood serum specimen, wherein the second blood serum specimen is an outgroup comprising cells from at least one mammal not afflicted with the disease. c) determining m/z values by mass spectrometry for proteins contained in the diseased specimen and the outgroup; e) comparing the m/z values of the outgroup with those of the diseased specimen and providing a score to m/z values in the outgroup not corresponding to m/z values in the diseased specimen to produce a hierarchical cladogram; and f) interpreting the cladogram.
7 . The method according to claim 6 , wherein the cladogram comprises a topographical hierarchy across a disease continuum, and wherein the disease continuum comprises a non-cancerous clade, a transitional non-cancerous clade, and a cancerous clade.
8 . The method according to claim 7 , wherein the cancerous clade further comprises a terminal clade and a middle clade, the terminal and middle clade derived from unique apomorphic protein changes in the diseased specimen.
9 . The method according to claim 8 , wherein the second blood serum specimen is derived from at least 50 humans.
10 . The method according to claim 8 , wherein the second blood serum specimen is derived from at least 50 primates.
11 . The method according to claim 10 wherein the at least 50 primates are chimpanzees.
12 . A method of treating cancer comprising:
a) determining m/z values of proteins from a patient afflicted with a cancer type by running the patient's blood plasma through a mass spectrometer; b) using a computer program to compare the m/z values of proteins in the cancer patient's blood plasma against a computer database containing a cladogram comprising m/z values of blood plasma derived from mammals having the cancer type, wherein the cladogram reflects a cumulative gradient of derived states comprising a terminal cancerous clade and a lower end cancerous clade, and wherein the terminal cancerous clade and the and the lower end cancerous clade are delineated from MS peptide peaks by a computer program; c) assessing a score to the patient's health condition; and d) assigning a stage to the cancer based on the patient's blood plasma derived m/z values on the cladogram and the assessed health condition score.
13 . The method according to claim 12 further comprising using a computer program to identify the peptides that make up the terminal cancerous clade and the lower end cancerous clade.
14 . A method of determining the progressive state of disease in an individual comprising:
a) obtaining a first gene expression profile for at least one cell, wherein the at least one cell is derived from a site of a disease in a patient; b) obtaining at least one additional gene expression profile for at least one additional cell, wherein the at least one additional cell is derived from a healthy tissue, bone, or plasma where the disease occurs; c) using a computer program to analyze the first gene expression profile and the at least one additional gene expression profile, wherein the computer program determines whether each gene in the first gene expression profile has an expression level outside of an expression level of a corresponding gene in the at least one additional gene expression profile, and provides a score to each gene in the first gene expression profile; d) using a computer to assess the score of each gene in the first gene expression profile and generate a parsimonious phylogenetic analyses from its assessment.
15 . The method according to claim 14 further comprising, using a computer to generate a cladogram, wherein the cladogram comprises a topographical hierarchy across a disease continuum, and wherein the disease continuum comprises a non-diseased clade, a transitional non-diseased clade, and a diseased clade.
16 . The method according to claim 14 , wherein the computer program determines:
a) whether each gene in the first gene expression profile has an expression level that is higher than the expression level of a corresponding gene in the at least one additional gene expression profile; b) whether each gene in the first gene expression profile has an expression level that is less than the expression level of a corresponding gene in the at least one additional gene expression profile; and c) whether each gene in the first gene expression profile has an expression level that is higher than and less than the expression level of a corresponding gene in the at least one additional gene expression profile.
17 . The method according to claim 14 , wherein the score is a weighted score, the weighted score varying between data points to emphasize or de-emphasize particular values.
18 . The method according to claim 17 , wherein the weighted score has a value between 1 and 0.
19 . The method according to claim 14 , wherein the disease is a type of cancer.
20 . The method according to claim 19 , wherein the type of cancer is selected from the group consisting of: breast, colon, lung, liver, ovarian, prostate, and pancreatic.
21 . The method according to claim 14 , wherein the at least one additional gene expression profile is obtained from at least one gene expression dataset.
22 . The method according to claim 14 , wherein the score is a first value or a second value, the first value assessed to genes having an expression level outside of an expression level of a corresponding gene in the at least one additional gene expression profile, and the second value assessed to genes having an expression level within an expression level of a corresponding gene in the at least one additional gene expression profile.
23 . The method according to claim 22 , wherein the first value is 1 and the second value is 0.Join the waitlist — get patent alerts
Track US2018188228A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.