US2021363597A1PendingUtilityA1
Identification and use of circulating nucleic acids
Assignee: UNIV LELAND STANFORD JUNIORPriority: Sep 12, 2014Filed: Aug 2, 2021Published: Nov 25, 2021
Est. expirySep 12, 2034(~8.1 yrs left)· nominal 20-yr term from priority
C40B 40/06C12N 15/1065C12Q 1/6806C12Q 1/6855C12N 15/11C12Q 1/6874C12Q 1/6886C40B 20/04C12Q 2600/156G16B 20/20
54
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Disclosed herein are polynucleotide adaptors and methods of use thereof for identifying and analyzing nucleic adds, including cell-free nucleic acids from a patient sample. Also disclosed herein are methods of using the adaptors to detect, diagnose, or determine prognosis of cancers.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A pool of unique adaptors for analyzing nucleic acids in a sample, each adaptor comprising:
a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single-stranded portion comprises:
i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and
ii) a random single-stranded barcode of at least two nucleotides specific to the adaptor.
2 . The pool of adaptors of claim 1 , wherein the double-stranded portion further comprises one or more G/C base pairs between the double-stranded barcode of at least two base pairs and the proximal end of the adaptor.
3 . The pool of adaptors of claim 2 , wherein the number of G/C base pairs varies among the adaptors in the pool.
4 . The pool of adaptors of claim 1 , wherein the double-stranded barcode comprises 2-20 base pairs.
5 . The pool of adaptors of claim 1 , wherein the pre-defined single-stranded barcode comprises 4-20 nucleotides.
6 . The pool of adaptors of claim 1 , wherein the random single-stranded barcode comprises 4-20 nucleotides.
7 . A method of analyzing nucleic acids comprising:
a) attaching a pool of adaptors according to claims 1 - 6 to both ends of a plurality of double-stranded nucleic acids via the double stranded portions of the adaptors, b) amplifying both strands of the adaptor-nucleic acids to produce first amplicons and second amplicons, wherein the first amplicons are derived from a first strand of the double-stranded nucleic acids and contain a first strand of the double-stranded barcodes, and the second amplicons are derived from a second strand of the double-stranded nucleic acids and contain a second strand of the double-stranded barcodes; c) determining the sequence of the first and second amplicons; and d) determining whether the first and the second amplicons originate from a single double-stranded nucleic acid of the plurality of the double-stranded nucleic acids by means of identifying the double-stranded barcode.
8 . The method of claim 7 , wherein the plurality of double-stranded nucleic acids comprises cell-free DNAs.
9 . The method of claim 7 , wherein the amplifying comprises 12-14 cycles of PCR.
10 . A method of analyzing a plurality of double-stranded nucleic acids, the method comprising:
a) attaching a pool of adaptors according to claims 1 - 6 to both ends of the plurality of double-stranded nucleic acids, b) amplifying both strands of the adaptor-nucleic acids to produce first amplicons and second amplicons, wherein the first amplicons are derived from a first strand of the double-stranded nucleic acids and contain a first strand of the double-stranded barcodes, and the second amplicons are derived from a second strand of the double-stranded nucleic acids and contain a second strand of the double-stranded barcodes; c) determining the sequence of the first and second amplicons; and d) identifying mutations in the first and second amplicon, wherein the mutation from the first and second amplicon are consistent mutations; or e) eliminating mutations that occur in the first but not the second amplicon; or f) eliminating G to T mutations that occur on at least about 90% of first amplicons derived from a first strand of a double-stranded nucleic acid, wherein the G to T mutations do not occur on less than about 10% of second amplicons derived from a second strand of the double-stranded nucleic acid; or g) eliminating mutations that are less than 100 base pairs from one another; or h) eliminating mutations that occur on less than about 50% of amplicons comprising the same pre-defined single stranded barcode and random single-stranded barcode; or i) any combination thereof.
11 . The method of claim 10 , wherein the first amplicons and the second amplicons of c) comprise the same endogenous barcode and the same double-stranded barcode, and wherein the first amplicons and the second amplicons of c) comprise different random barcodes derived from the random single-stranded barcode of the adaptor.
12 . The method of claim 10 , wherein g) comprises eliminating mutations that are less than 5 base pairs from another.
13 . The method of claim 10 , wherein h) comprises eliminating mutations that occur on less than about 60%, about 70%, about 80%, about 90%, about 95%, or about 100% of amplicons comprising the same double-stranded stem barcode and the same endogenous barcode.
14 . A method of reduced-error analysis of nucleic acid comprising:
a) attaching to each end of nucleic acids an adaptor from a pool of unique adaptors each adaptor comprising a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single stranded portion containing a 5′-terminal nucleotide comprises: i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and ii) a random single-stranded barcode of at least two nucleotides specific to one strand of the adaptor; b) sequencing the nucleic acids with attached adaptors to determine sequence and if present, sequence variations of the nucleic acids; c) grouping the sequences of nucleic acids sharing the same random single-stranded barcode specific to one strand of the adaptor, to form barcode groups; d) eliminating sequence variations that are present in fewer than all members of the barcode group; e) eliminating sequence variations that are present at a frequency below a predetermined threshold among the barcode groups.
15 . The method of claim 14 , wherein the predetermined threshold is 50%.
16 . The method of claim 14 , wherein the threshold is predetermined according to a method comprising the steps of:
a) performing single molecule sequencing of multiple samples to determine the target nucleic acid sequence; b) for each of the possible classes of nucleotide substitutions, determining a total number of substitutions (y) in all positions; and a number of supporting reads (t) for each position having a substitution; a) defining a function relating y to t; d) solving the function for the desired value of y by determining t, wherein t is the threshold number of reads above which the substitution may be called a sequence variation at the base position in the nucleic acid.
17 . A method of analyzing nucleic acids in a sample comprising:
a) attaching to each end of nucleic acids an adaptor from a pool of unique adaptors each adaptor comprising a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single stranded portion containing a 5′ terminal nucleotide comprises: i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and ii) a random single-stranded barcode of at least two nucleotides specific to one strand of the adaptor; b) sequencing the nucleic acids with attached adaptors to determine sequence and if present, sequence variations of the nucleic acids; c) grouping the sequences of nucleic acids sharing the same random single-stranded barcode to form barcode groups; d) eliminating sequence variations that are present in fewer than all members of a barcode group; e) performing steps a)-d) on nucleic acids from control samples to identify recurrent sequence variations; f) applying statistical analysis to determine a confidence interval for the frequency of each sequence variation identified in step e); g) setting a threshold for the frequency of sequence variations within the confidence interval of step f); h) eliminating sequence variations whose frequency falls below the threshold set in step g).
18 . A method of assessing a patient by analyzing patient's cell-free nucleic acids by the method of claim 17 , further comprising step i) of assessing the patient as having cancer if one or more of the sequence variations not eliminated in steps d) and h) are present.
19 . A method of designing a selector comprising a plurality of target genomic regions to be analyzed in a sample of a patient having a type of tumor, the method comprising:
a) performing sequencing of a genome of the type of tumor from multiple patients; b) identifying regions of the genome containing a mutation; c) ranking the regions identified in step b) based on the highest number of patients having a mutation per kilobase of sequence obtained in step a); d) ranking the regions identified in step b) based on the highest number of patients having a mutation per exon sequenced in step a); e) including the highest ranked regions from steps c) and d) in the selector.
20 . The method of claim 19 , wherein the genome sequencing in step a) is exon sequencing.
21 . The method of claim 19 , wherein regions identified in step b) are at least 100 base pairs long.
22 . The method of claim 19 , wherein the mutations comprise single nucleotide variations, copy number variations, fusions, seed regions and histology classification regions.
23 . The method of claim 19 , wherein the highest ranked regions included in the selector comprise the top 10% of the highest ranking regions.
24 . The method of claim 19 , further comprising after step b), eliminating from the selector regions that fall into repeat-rich regions of the genome.
25 . A method of assessing cancer in a patient comprising:
a) designing a selector according to claim 19 ; b) obtaining a sample from a patient comprising cell-free nucleic acids; c) determining the sequence of genomic regions of the selector in the patient's nucleic acids; d) assessing the patient as likely to have cancer or recurrence of cancer if at least one sequence determined in step c) contains a mutation.
26 . The method of claim 25 , further comprising a confirmation of mutations detected in step b) as somatic in a matched tumor biopsy.
27 . A method of setting a threshold for calling a sequence variant at a base position in a target nucleic acid sequence containing nucleotide substitutions, the method comprising:
a) performing single molecule sequencing of barcoded nucleic acids from multiple samples to determine the target nucleic acid sequence; b) for each of the possible classes nucleotide substitutions, determining a total number of substitutions (y) in all positions; a number of supporting reads (t) for the position having a substitution; c) defining a function relating y to t; d) solving the function for the desired value of y by determining t, wherein t is the threshold number of reads above which the substitution may be called a variant at the base position in the nucleic acid.
28 . The method of claim 27 , wherein the threshold t for a given sequence g among the plurality of target sequences is adjusted for global error rate by a method comprising the steps of:
a) determining error rate e for the plurality of target sequences equal to the number of base positions with nucleotide substitutions in a target sequence divided by the total number of bases in the target sequence; b) determining sequencing depth d for the plurality of target sequences; c) if e for sequence g falls within the top 25% of e of the plurality of target sequences, the threshold t for sequence g is adjusted to t′ according to the formula: t′←t×w, where w=min{q 2 , 5} and q=e divided by the 75t h percentile of the error rates of sequences in the selector; d) if d for sequence g falls below the median of sequencing depths of the plurality of target sequences (d med ), the threshold t for sequence g is adjusted to t′ according to the formula: t′←t/w*, where w*=ln(d med /d);
29 . A method of assessing a non-small cell lung cancer (NSCLC) patient by analyzing the patient's cfDNA according to claim 17 , further comprising step i) assessing the patient as assessing the patient as having NSCLC or having a progression of NSCLC if one or more of the sequence variations not eliminated in steps d) and h) are present.
30 . The method of claim 29 , wherein the mutation is a mutation in epidermal growth factor receptor (EGFR) gene located in the kinase domain (exon 19, 20 and 21) of the gene.
31 . A method of pairing nucleic acid sequencing reads to obtain a double-stranded nucleic acid sequence comprising:
a) determining the sequence of plurality of single-stranded nucleic comprising insert sequences and adaptor sequences containing barcodes; b) determining genomic coordinates of the insert sequences; c) pairing the sequences into a double-stranded nucleic acid if the sequences have complementary barcodes and genomic coordinates of the insert map to the opposite strands.
32 . The method of claim, further comprising a step of eliminating single-member barcode families containing a sequence variant unless the variant is supported by at least one other barcode family with members.
33 . A pool of unique adaptors for analyzing nucleic acids in a sample, each adaptor comprising:
a double stranded portion at a proximal end and at least one single stranded portion at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single-stranded portion comprises:
i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and
ii) a random single-stranded barcode of at least two nucleotides specific to the adaptor.
34 . The pool of unique adaptors of claim 33 , each adaptor comprising two single-stranded portions at the distal end; one portion comprising a 5′-end and the other portion comprising a 3′-end, wherein the single stranded portions are non-hybridizable with each other.
35 . The pool of unique adaptors of claim 34 , wherein the two single stranded portions are covalently linked to each other at the distal ends.
36 . The pool of unique adaptors of claim 35 , wherein the two single stranded portions are covalently linked to each other via a linker.
37 . The pool of unique adaptors of claim 36 , wherein the linker comprises a cleavage site.
38 . The pool of unique adaptors of claim 33 , comprising a combination of two sub-pools of adaptors:
i) a first sub-pool wherein each adaptor comprises two single-stranded portions at the distal end: one portion comprising a 5′-end and the other portion comprising a 3′-end, wherein the single stranded portions are non-hybridizable with each other; and ii) a second sub-pool wherein each adaptor comprises two non-hybridizable single-stranded portions that are covalently linked to each other at the distal ends.
39 . A method of reduced-error analysis of nucleic acid in a subject's sample comprising:
a) performing single molecule sequencing nucleic acids from multiple control samples to determine the target nucleic acid sequence; b) determining the frequency of each of the possible classes of nucleotide substitutions at each position among the control samples; c) fitting a statistical model to the frequencies determined in step b) to determine frequencies of background errors; d) performing single molecule sequencing nucleic acids from the subject's sample; e) determining the frequency of each of the possible classes of nucleotide substitutions at each position in the subject's sample; f) determining the depth of reads for each target sequence in the subject's sample; g) applying the statistical model from step c) to the frequencies and depth determined in steps e) and f); h) eliminating nucleotide substitutions having frequencies below frequencies of background errors determined in step c).Join the waitlist — get patent alerts
Track US2021363597A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.