US2016032396A1PendingUtilityA1
Identification and Use of Circulating Nucleic Acid Tumor Markers
Assignee: UNIV LELAND STANFORD JUNIORPriority: Mar 15, 2013Filed: Mar 12, 2014Published: Feb 4, 2016
Est. expiryMar 15, 2033(~6.6 yrs left)· nominal 20-yr term from priority
C12Q 1/6886G06F 19/22C12Q 2600/156G16B 30/10C12Q 1/6806C12Q 1/6827G16B 30/00C12Q 1/6855
69
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods for creating a selector of mutated genomic regions and for using the selector set to analyze genetic alterations in a cell-free nucleic acid sample are provided. The methods can be used to measure tumor-derived nucleic acids in a blood sample from a subject and thus to monitor the progression of disease in the subject. The methods can also be used for cancer screening, cancer diagnosis, cancer prognosis, and cancer therapy designation.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of detecting, diagnosing, prognosing, or therapy selection of a cancer in a subject in need thereof, the method comprising:
(a) obtaining sequence information of a cell-free DNA (cfDNA) sample derived from the subject; and (b) using the sequence information derived from (a) to detect circulating tumor DNA (ctDNA) in the sample, wherein the method is capable of detecting a percentage of ctDNA that is less than or equal to 2% of total cfDNA.
2 . The method of claim 1 , wherein the method is capable of detecting a percentage of ctDNA that is less than or equal to 1.75%, 1.5%, 1.25%, 1%, 0.75%, 0.50%, 0.25%, 0.1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.05%, 0.01%, 0.009%, 0.008%, 0.007%, 0.006%, 0.005%, 0.004%, 0.003%, 0.002%, 0.001%, 0.0005%, or 0.00001% of the total cfDNA.
3 . The method of claim 1 , wherein the sample is a plasma, serum, sweat, breath, tears, saliva, urine, stool, amniotic fluid, or cerebral spinal fluid sample.
4 . The method of claim 1 , wherein the sample is not a pap smear, cyst fluid, or pancreatic fluid sample.
5 . The method of claim 1 , wherein the sequence information comprises information related to at least 2, 3, 5, 8, 10, 20, 30, 40, 100, 200, or 300 genomic regions.
6 . The method of claim 5 , wherein the genomic regions comprise two or more of exonic regions, intronic regions, and untranslated regions.
7 . The method of claim 5 , wherein the genomic regions comprise less than 1.5 megabases (Mb), 1 Mb, 500 kb, 350 kb, 100 kb, 75 kb, 50 kb or 25 kb of the genome.
8 . The method of claim 1 , wherein the sequence information comprises information pertaining to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 or more genomic regions from a selector set comprising a plurality of genomic regions.
9 . The method of claim 8 , wherein the plurality of genomic regions are based on a selector set comprising genomic regions comprising one or more mutations present in one or more subjects from a population of cancer subjects.
10 . The method of claim 8 , wherein at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the plurality of genomic regions are based on a selector set comprising genomic regions comprising one or more mutations present in one or more subjects from a population of cancer subjects.
11 . The method of claim 9 or 10 , wherein the selector set comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100 or more genomic regions selected from any one of Tables 2 and 18.
12 . The method of claim 1 , wherein the obtaining sequence information of step (a) comprises performing massively parallel sequencing.
13 . The method of claim 1 , wherein the obtaining sequence information of step (a) comprises using one or more adaptors.
14 . The method of claim 13 , wherein the one or more adaptors comprise a molecular barcode comprising a randomer sequence.
15 . The method of claim 1 , wherein using the sequence information of step (b) comprises detecting one or more of SNVs, indels, copy number variants, and rearrangements in selected regions of the subject's genome.
16 . The method of claim 1 , wherein using the sequence information of step (b) comprises detecting two or more of SNVs, indels, copy number variants, and rearrangements in selected regions of the subject's genome.
17 . The method of claim 1 , wherein the detecting of step (b) does not involve performing digital PCR (dPCR).
18 . The method of claim 1 , wherein the detecting of step (b) comprises applying an algorithm to the sequence information to determine a quantity of one or more genomic regions from a selector set.
19 . The method of claim 1 , further comprising detecting, diagnosing, prognosing or selecting a therapy for a cancer in the subject based on the detection of ctDNA.
20 . The method of claim 19 , wherein diagnosing or prognosing the cancer has a sensitivity of at least about 50%, 52%, 55%, 57%, 60%, 62%, 65%, 67%, 70%, 72%, 75%, 77%, 80%, 82%, 85%, 87%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%.
21 . The method of claim 19 , wherein diagnosing or prognosing the cancer has a specificity of at least about 50%, 52%, 55%, 57%, 60%, 62%, 65%, 67%, 70%, 72%, 75%, 77%, 80%, 82%, 85%, 87%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%.
22 . A method of producing a selector set for a cancer comprising:
(a) identifying genomic regions comprising mutations in one or more subjects from a population of subjects suffering from the cancer; (b) ranking the genomic regions based on a Recurrence Index (RI), wherein the RI of the genomic region is determined by dividing the number of subjects or tumors with mutations in the genomic region by the size of the genomic region; and (c) producing a selector set based on the RI.
23 . The method of claim 22 , wherein at least a subset of the genomic regions are exon regions, intron regions, untranslated regions, or a combination thereof.
24 . The method of claim 22 , wherein producing the selector set based on the RI comprises selecting genomic regions that have a recurrence index in the top 70 th , 75 th , 80 th , 85 th , 90 th , or 95 th or greater percentile.
25 . The method of claim 22 , wherein producing the selector set comprises applying an algorithm to a subset of the ranked genomic regions.
26 . The method of claim 22 , wherein producing the selector set comprises selecting genomic regions that maximize a median number of mutations per subject of the selector set.
27 . The method of claim 22 , wherein producing the selector set comprises selecting genomic regions that maximize the number of subjects in the selector set.
28 . The method of claim 22 , wherein producing the selector set comprises selecting genomic regions that minimize the total size of the genomic regions.
29 . A computer readable medium comprising sequence information for two or more genomic regions wherein:
(a) the two or more genomic regions comprise one or more mutations present in greater than or equal to 80% of tumors from a first population of subjects suffering from a first type of cancer; (b) the two or more genomic regions represent less than 1.5 Mb of the genome; and (c) one or more of the following:
(i) the condition is not hairy cell leukemia, ovarian cancer, Waldenstrom's macroglobulinemia;
(ii) a genomic region comprises at least one mutation in at least one subject afflicted with the cancer;
(iii) the two or more genomic regions comprise one or more mutations present in a second population of subjects suffering from a second type of cancer;
(iv) the two or more genomic regions are derived from two or more different genes;
(v) the genomic regions comprise two or more mutations; or
(vi) the two or more genomic regions comprise at least 10 kb.
30 . The computer readable medium of claim 29 , wherein the genomic regions comprise one or more mutations present in greater than or equal to 60% of tumors from the second population of subjects suffering from the second type of cancer.
31 . The computer readable medium of claim 29 , wherein the genomic regions are derived from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more different genes.
32 . The computer readable medium of claim 29 , wherein the genomic regions comprise at least 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 kb.
33 . The computer readable medium of claim 29 , wherein the sequence information comprises genomic coordinates pertaining to the two or more genomic regions.
34 . The computer readable medium of claim 29 , wherein the sequence information comprises a nucleic acid sequence pertaining to the two or more genomic regions.
35 . The computer readable medium of claim 29 , wherein the sequence information comprises a length of the two or more genomic regions.
36 . A composition comprising a set of oligonucleotides that selectively hybridize to a plurality of genomic regions, wherein:
(a) greater than or equal to 80% of tumors from a population of cancer subjects include one or more mutations in the genomic regions; (b) the plurality of genomic regions represent less than 1.5 Mb of the genome; and (c) the set of oligonucleotides comprise 5 or more different oligonucleotides that selectively hybridize to the plurality of genomic DNA regions.
37 . The composition of claim 36 , wherein the genomic DNA regions comprise at least 2 regions from those identified in any one of Tables 2 and 6-18.
38 . The composition of claim 36 , wherein the set of oligonucleotides hybridize to between about 5 kb to 1000 kb of the genome.
39 . The composition of claim 36 , wherein the set of oligonucleotides are capable of hybridizing to 5 or more different genomic regions.
40 . The composition of claim 36 , wherein the oligonucleotides are attached to a solid support.
41 . The composition of claim 40 , wherein the solid support is a bead.
42 . The composition of claim 40 , wherein the solid support is an array.
43 . A method for preparing a library for sequencing comprising:
(a) conducting an amplification reaction on cell-free DNA (cfDNA) derived from a sample to produce a plurality of amplicons, wherein the amplification reaction comprises 20 or fewer amplification cycles; and (b) producing a library for sequencing, the library comprising the plurality of amplicons.
44 . The method of claim 43 , wherein the amplification reaction comprises 15 or fewer amplification cycles.
45 . The method of claim 43 , further comprising attaching adaptors to the cell-free DNA.
46 . The method of claim 45 , wherein the adaptors comprise a molecular barcode.
47 . The method of claim 45 , wherein the adaptors comprise a sample index.
48 . The method of claim 45 , wherein the adaptors comprise a primer sequence.
49 . The method of claim 45 , wherein the adaptors comprise a Y-shaped adaptor.
50 . The method of claim 43 , further comprising fragmenting the cfDNA.
51 . The method of claim 43 , further comprising end-repairing the cfDNA.
52 . The method of claim 43 , further comprising A-tailing the cfDNA.
53 . A method of determining a statistical significance of a selector set, the method comprising:
(a) detecting a presence of one or more mutations in one or more samples from a subject, wherein the one or more mutations are based on a selector set comprising genomic regions comprising the one or more mutations; (b) determining a mutation type of the one or more mutations present in the sample; and (c) determining a statistical significance of the selector set by calculating a ctDNA detection index based on a p-value of the mutation type of mutations present in the one or more samples.
54 . The method of claim 53 , wherein if a rearrangement is observed in two or more samples from the subject, then the ctDNA detection index is 0.
55 . The method of claim 54 , wherein at least one of the two or more samples is a plasma sample.
56 . The method of claim 54 , wherein at least one of the two or more samples is a tumor sample.
57 . The method of claim 54 , wherein the rearrangement is a fusion or a breakpoint.
58 . The method of claim 53 , wherein if one type of mutation is present, then the ctDNA detection index is the p-value of the one type of mutation.
59 . The method of claim 53 , wherein if: (i) two or more types of mutations are present in the sample; (ii) the p-values of the two or more types mutations are less than 0.1; and (iii) a rearrangement is not one of the types of mutations, then the ctDNA detection is calculated based on the combined p-values of the two or more mutations.
60 . The method of claim 59 , wherein the p-values of the two or more mutations are combined according to Fisher's method.
61 . The method of claim 59 , wherein one of the two or more types of mutations is a SNV.
62 . The method of claim 61 , wherein the p-value of the SNV is determined by Monte Carlo sampling.
63 . The method of claim 59 , wherein one of the two or more types of mutations is an indel.
64 . The method of claim 53 , wherein if: (i) two or more types of mutations are present in the sample; (ii) a p-value of at least one of the two or more types of mutations are greater than 0.1; and (iii) a rearrangement is not one of the types of mutations, then the ctDNA detection is calculated based on the p-value of one of the two or more types mutations.
65 . The method of claim 64 , wherein one of the two or more types of mutations is a SNV.
66 . The method of claim 65 , wherein the ctDNA detection index is calculated based on the p-value of the SNV.
67 . The method of claim 64 , wherein one of the two or more types of mutations is an indel.
68 . A method of identifying rearrangements in one or more nucleic acids, the method comprising:
(a) obtaining sequencing information pertaining to a plurality of genomic regions; (b) producing a list of genomic regions, wherein the genomic regions are adjacent to one or more candidate rearrangement sites or the genomic regions comprise one or more candidate rearrangement sites; (c) applying an algorithm to the list of genomic regions to validate candidate rearrangement sites, thereby identifying rearrangements.
69 . The method of claim 68 , wherein the sequencing information comprises an alignment file.
70 . The method of claim 69 , wherein the alignment file comprises an alignment file of pair-end reads, exon coordinates, and a reference genome.
71 . The method of claim 68 , wherein the sequencing information is obtained from a database.
72 . The method of claim 68 , wherein the sequencing information is obtained from one or more samples from one or more subjects.
73 . The method of claim 68 , wherein producing the list of genomic regions comprises identifying discordant read pairs based on the sequencing information.
74 . The method of claim 73 , wherein producing the list of genomic regions comprises classifying the discordant read pairs based on the sequencing information.
75 . The method of claim 73 , wherein producing the list of genomic regions further comprises ranking the genomic regions.
76 . The method of claim 75 , wherein the genomic regions are ranked in decreasing order of discordant read depth.
77 . The method of claim 68 , wherein producing the list of genomic regions comprises using an algorithm to analyze properly paired reads in which one of the paired reads is truncated to produce a soft-clipped read.
78 . The method of claim 68 , wherein the algorithm analyzes the soft-clipped reads based on a pattern.
79 . The method of claim 78 , wherein the pattern is based on x number of skipped bases (Sx) and on y number of contiguous mapped bases (My).
80 . The method of claim 79 , wherein the pattern is MySx or SxMy.
81 . The method of claim 68 , wherein applying the algorithm to validate the candidate rearrangement sites comprises ranking the candidate rearrangements based on their read frequency.
82 . The method of claim 68 , wherein applying the algorithm to validate the candidate rearrangement sites comprises comparing two or more reads of the candidate rearrangement.
83 . The method of claim 82 , wherein applying the algorithm to validate the candidate rearrangement sites comprises identifying the candidate rearrangement as a rearrangement if the two or more reads have a sequence alignment.
84 . A method of identifying tumor-derived single nucleotide variations (SNVs), the method comprising:
(a) obtaining a sample from a subject suffering from a cancer or suspected of suffering from a cancer; (b) conducting a sequencing reaction on the sample to produce sequencing information; (c) applying an algorithm to the sequencing information to produce a list of candidate tumor alleles based on the sequencing information from step (b), wherein a candidate tumor allele comprises a non-dominant base that is not a germline SNP; and (d) identifying tumor-derived SNVs based on the list of candidate tumor alleles.
85 . The method of claim 84 , wherein producing the list of candidate tumor alleles comprises ranking the tumor alleles by their fractional abundance.
86 . The method of claim 85 , wherein producing the list of candidate tumor alleles comprises ranking the tumor alleles based on a sequencing depth.
87 . The method of claim 86 , wherein producing the list of candidate tumor alleles comprises selecting tumor alleles that meet a minimum sequencing depth.
88 . The method of claim 87 , wherein the minimum sequencing depth is at least 100×, 200×, 300×, 400×, 500×, 600×, 700×, 800×, 900×, 1000× or more.
89 . A method of producing a selector set comprising:
(a) obtaining sequencing information of a tumor sample from a subject suffering from a cancer; (b) comparing the sequencing information of the tumor sample to sequencing information from a non-tumor sample from the subject to identify one or more mutations specific to the sequencing information of the tumor sample; and (c) producing a selector set comprising one or more genomic regions comprising the one or more mutations specific to the sequencing information of the tumor sample.
90 . The method of claim 89 , wherein the selector set comprises sequencing information pertaining to the one or more genomic regions.
91 . The method of claim 90 , wherein the selector set comprises genomic coordinates pertaining to the one or more genomic regions.
92 . The method of claim 90 , wherein the selector set comprises a plurality of oligonucleotides that selectively hybridize the one or more genomic regions.
93 . The method of claim 92 , wherein the plurality of oligonucleotides are biotinylated.
94 . The method of claim 89 , the one or more mutations comprise SNVs, indels, rearrangements, or a combination thereof.
95 . The method of claim 94 , wherein producing the selector set comprises identifying tumor-derived SNVs based on the method of any one of claims 84 - 88 .
96 . The method of claim 94 , wherein producing the selector set comprises identifying tumor-derived rearrangements based on the method of any one of claims 68 - 83 .Join the waitlist — get patent alerts
Track US2016032396A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.