US2016153054A1PendingUtilityA1
Biomarkers for colorectal cancer
Est. expiryAug 6, 2033(~7 yrs left)· nominal 20-yr term from priority
C12Q 2600/158C12Q 2600/16C12Q 1/6886C12Q 1/6806C12Q 2600/112
53
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Biomarkers and methods for predicting the risk of a disease related to microbiota, in particular colorectal cancer (CRC), are described.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of obtaining a set of gene markers for predicting the risk of an abnormal condition related to microbiota, comprising:
a) identifying abnormal-associated gene markers by a metagenome-wide association study (MGWAS) strategy comprising:
i) collecting a sample from each subject from a population of subjects with the abnormal condition (abnormal) and subjects without the abnormal condition (controls),
ii) extracting DNA from each sample, constructing a DNA library from each sample, and carrying out high-throughput sequencing of each DNA library to obtain sequencing reads for each sample;
iii) mapping the sequencing reads to a gene catalog, and deriving a gene profile from the mapping result;
iv) performing a Wilcoxon rank-sum test on the gene profile to identify differential metagenomic gene contents between the abnormal and controls;
b) ranking all of the abnormal-associated gene markers identified in step a) by a minimum redundancy-maximum relevance (mRMR) method, and identifying/classifying sequential marker sets therefrom; and c) for each sequential marker set, estimating the error rate by leave-one-out cross-validation (LOOCV) of a linear discrimination classifier, and selecting an optimal gene marker set with the lowest error rate as the set of gene markers for predicting the risk of the abnormal condition.
2 . A method of diagnosing whether a subject has an abnormal condition related to microbiota or is at the risk of developing an abnormal condition related to microbiota, comprising:
1) obtaining sequencing reads from sample j of the subject; 2) mapping the sequencing reads to a gene catalog and deriving a gene profile from the mapping result; 3) determining the relative abundance of each gene marker in a set of gene markers, wherein the set of gene markers is obtained using the method of claim 1 ; and 4) calculating the index of sample j using the following formula:
I
j
=
[
∑
i
ε
N
log
10
(
A
ij
+
10
-
20
)
N
-
∑
i
ε
M
log
10
(
A
ij
+
10
-
20
)
M
]
,
wherein:
A ij is the relative abundance of marker i in sample j, wherein i refers to each of the gene markers in the gene marker set,
N is a subset of all of the abnormal-associated gene markers in selected biomarkers related to the abnormal condition,
M is a subset of all of the control-associated gene markers in selected biomarkers related to the abnormal condition, and
|N| and |M| are numbers (sizes) of the biomarkers in these two subsets, respectively, wherein an index greater than a cutoff indicates that the subject has or is at the risk of developing the abnormal condition.
3 . The method of claim 1 , wherein the metagenome-wide association study (MGWAS) strategy further comprises estimating the false discovery rate (FDR).
4 . The method of claim 2 , wherein the gene catalog is a non-redundant gene set constructed for related microbiota.
5 . The method of claim 2 , wherein the abnormal condition related to microbiota is an abnormal condition related to environmental microbiota such as soil microbiota, sea microbiota, or river microbiota.
6 . The method of claim 2 , wherein the abnormal condition related to microbiota is a disease related to microbiota present in the animal body or the human body, wherein the microbiota is selected from the group consisting of microbiota found in the gastrointestinal tract, nasal passages, oral cavities, skin and the urogenital tract.
7 . The method of claim 2 , wherein the abnormal condition related to microbiota is a colorectal disease selected from the group consisting of Colorectal Cancer, Ulcerative Colitis, Crohn's Disease, Irritable Bowel Syndrome (IBS), Diverticular Disease, Hemorrhoids, Anal Fissure, and Bowel Incontinence.
8 . The method of claim 2 , wherein the sequencing reads are obtained via a next-generation sequencing method or a next-next-generation sequencing method.
9 . The method of claim 8 , wherein the sequencing reads are obtained via at least one system selected from the group consisting of Hiseq 2000, SOLID, 454, and True Single Molecule Sequencing.
10 . The method of claim 2 , wherein the cutoff value is obtained by a Receiver Operator Characteristic (ROC) method, wherein the cutoff corresponds to the value when the AUC (Area Under the Curve) is at its maximum.
11 . A method for diagnosing whether a subject has colorectal cancer (CRC) or is at the risk of developing colorectal cancer, comprising:
1) obtaining sequencing reads from sample j of the subject; 2) mapping the sequencing reads to a human gut gene catalog and deriving a gene profile from the mapping result; 3) determining the relative abundance of each of the gene markers listed in SEQ ID NOs: 1-31; and 4) calculating the index of sample j using the following formula:
I
j
=
[
∑
i
ε
N
log
10
(
A
ij
+
10
-
20
)
N
-
∑
i
ε
M
log
10
(
A
ij
+
10
-
20
)
M
]
,
wherein:
A ij is the relative abundance of marker i in sample j, wherein i refers to each of the gene markers listed in SEQ ID NOs 1-31,
N is a subset of all of CRC-associated gene markers and M is a subset of all of the control-associated gene markers,
wherein the subset of CRC-associated gene markers and the subset of control-associated gene markers are shown in Table 1, and
|N| and |M| are numbers (sizes) of the biomarkers in these two subsets, respectively,
wherein an index greater than a cutoff indicates that the subject has or is at the risk of developing colorectal cancer.
12 . The method of claim 11 , wherein the cutoff value is obtained by a Receiver Operator Characteristic (ROC) method, wherein the cutoff corresponds to the value when the AUC (Area Under the Curve) is at its maximum.
13 . The method of claim 12 , wherein the value of the cutoff is −0.0575.
14 . A gene marker set for predicting the risk of colorectal cancer (CRC) in a subject, consisting of the genes listed in SEQ ID NOs: 1-31.
15 . A kit for analyzing the gene marker set of claim 14 , comprising primers used for PCR amplification that are designed according to the genes listed in SEQ ID NOs: 1-31.
16 . A kit for analyzing the gene marker set of claim 14 , comprising one or more probes that are designed according to the genes listed in SEQ ID NOs: 1-31.
17 . A method comprising using of the gene marker set of claim 14 for preparation of a kit for predicting the risk of colorectal cancer (CRC) in a subject.
18 . The method of claim 2 , wherein the sample is a feces sample, a nasal cavity swab, a buccal swab, a skin swab or a vaginal swab.
19 . The method of claim 2 , wherein the sequencing reads are obtained via steps comprising:
1) collecting the sample j from the subject and extracting DNA from the sample, and 2) constructing a DNA library and sequencing the library.
20 . The method of claim 11 , wherein the sequencing reads are obtained via steps comprising:
1) collecting the sample j from the subject and extracting DNA from the sample, 2) constructing a DNA library and sequencing the library.Join the waitlist — get patent alerts
Track US2016153054A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.