US2021363597A1PendingUtilityA1

Identification and use of circulating nucleic acids

Assignee: UNIV LELAND STANFORD JUNIORPriority: Sep 12, 2014Filed: Aug 2, 2021Published: Nov 25, 2021
Est. expirySep 12, 2034(~8.1 yrs left)· nominal 20-yr term from priority
C40B 40/06C12N 15/1065C12Q 1/6806C12Q 1/6855C12N 15/11C12Q 1/6874C12Q 1/6886C40B 20/04C12Q 2600/156G16B 20/20
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed herein are polynucleotide adaptors and methods of use thereof for identifying and analyzing nucleic adds, including cell-free nucleic acids from a patient sample. Also disclosed herein are methods of using the adaptors to detect, diagnose, or determine prognosis of cancers.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A pool of unique adaptors for analyzing nucleic acids in a sample, each adaptor comprising:
 a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single-stranded portion comprises:
 i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and 
 ii) a random single-stranded barcode of at least two nucleotides specific to the adaptor. 
   
     
     
         2 . The pool of adaptors of  claim 1 , wherein the double-stranded portion further comprises one or more G/C base pairs between the double-stranded barcode of at least two base pairs and the proximal end of the adaptor. 
     
     
         3 . The pool of adaptors of  claim 2 , wherein the number of G/C base pairs varies among the adaptors in the pool. 
     
     
         4 . The pool of adaptors of  claim 1 , wherein the double-stranded barcode comprises 2-20 base pairs. 
     
     
         5 . The pool of adaptors of  claim 1 , wherein the pre-defined single-stranded barcode comprises 4-20 nucleotides. 
     
     
         6 . The pool of adaptors of  claim 1 , wherein the random single-stranded barcode comprises 4-20 nucleotides. 
     
     
         7 . A method of analyzing nucleic acids comprising:
 a) attaching a pool of adaptors according to  claims 1 - 6  to both ends of a plurality of double-stranded nucleic acids via the double stranded portions of the adaptors,   b) amplifying both strands of the adaptor-nucleic acids to produce first amplicons and second amplicons, wherein the first amplicons are derived from a first strand of the double-stranded nucleic acids and contain a first strand of the double-stranded barcodes, and the second amplicons are derived from a second strand of the double-stranded nucleic acids and contain a second strand of the double-stranded barcodes;   c) determining the sequence of the first and second amplicons; and   d) determining whether the first and the second amplicons originate from a single double-stranded nucleic acid of the plurality of the double-stranded nucleic acids by means of identifying the double-stranded barcode.   
     
     
         8 . The method of  claim 7 , wherein the plurality of double-stranded nucleic acids comprises cell-free DNAs. 
     
     
         9 . The method of  claim 7 , wherein the amplifying comprises 12-14 cycles of PCR. 
     
     
         10 . A method of analyzing a plurality of double-stranded nucleic acids, the method comprising:
 a) attaching a pool of adaptors according to  claims 1 - 6  to both ends of the plurality of double-stranded nucleic acids,   b) amplifying both strands of the adaptor-nucleic acids to produce first amplicons and second amplicons, wherein the first amplicons are derived from a first strand of the double-stranded nucleic acids and contain a first strand of the double-stranded barcodes, and the second amplicons are derived from a second strand of the double-stranded nucleic acids and contain a second strand of the double-stranded barcodes;   c) determining the sequence of the first and second amplicons; and   d) identifying mutations in the first and second amplicon, wherein the mutation from the first and second amplicon are consistent mutations; or   e) eliminating mutations that occur in the first but not the second amplicon; or   f) eliminating G to T mutations that occur on at least about 90% of first amplicons derived from a first strand of a double-stranded nucleic acid, wherein the G to T mutations do not occur on less than about 10% of second amplicons derived from a second strand of the double-stranded nucleic acid; or   g) eliminating mutations that are less than 100 base pairs from one another; or   h) eliminating mutations that occur on less than about 50% of amplicons comprising the same pre-defined single stranded barcode and random single-stranded barcode; or   i) any combination thereof.   
     
     
         11 . The method of  claim 10 , wherein the first amplicons and the second amplicons of c) comprise the same endogenous barcode and the same double-stranded barcode, and wherein the first amplicons and the second amplicons of c) comprise different random barcodes derived from the random single-stranded barcode of the adaptor. 
     
     
         12 . The method of  claim 10 , wherein g) comprises eliminating mutations that are less than 5 base pairs from another. 
     
     
         13 . The method of  claim 10 , wherein h) comprises eliminating mutations that occur on less than about 60%, about 70%, about 80%, about 90%, about 95%, or about 100% of amplicons comprising the same double-stranded stem barcode and the same endogenous barcode. 
     
     
         14 . A method of reduced-error analysis of nucleic acid comprising:
 a) attaching to each end of nucleic acids an adaptor from a pool of unique adaptors each adaptor comprising a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single stranded portion containing a 5′-terminal nucleotide comprises: i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and ii) a random single-stranded barcode of at least two nucleotides specific to one strand of the adaptor;   b) sequencing the nucleic acids with attached adaptors to determine sequence and if present, sequence variations of the nucleic acids;   c) grouping the sequences of nucleic acids sharing the same random single-stranded barcode specific to one strand of the adaptor, to form barcode groups;   d) eliminating sequence variations that are present in fewer than all members of the barcode group;   e) eliminating sequence variations that are present at a frequency below a predetermined threshold among the barcode groups.   
     
     
         15 . The method of  claim 14 , wherein the predetermined threshold is 50%. 
     
     
         16 . The method of  claim 14 , wherein the threshold is predetermined according to a method comprising the steps of:
 a) performing single molecule sequencing of multiple samples to determine the target nucleic acid sequence;   b) for each of the possible classes of nucleotide substitutions, determining a total number of substitutions (y) in all positions; and a number of supporting reads (t) for each position having a substitution;   a) defining a function relating y to t;   d) solving the function for the desired value of y by determining t, wherein t is the threshold number of reads above which the substitution may be called a sequence variation at the base position in the nucleic acid.   
     
     
         17 . A method of analyzing nucleic acids in a sample comprising:
 a) attaching to each end of nucleic acids an adaptor from a pool of unique adaptors each adaptor comprising a double stranded portion at a proximal end and two single stranded portions at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single stranded portion containing a 5′ terminal nucleotide comprises: i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and ii) a random single-stranded barcode of at least two nucleotides specific to one strand of the adaptor;   b) sequencing the nucleic acids with attached adaptors to determine sequence and if present, sequence variations of the nucleic acids;   c) grouping the sequences of nucleic acids sharing the same random single-stranded barcode to form barcode groups;   d) eliminating sequence variations that are present in fewer than all members of a barcode group;   e) performing steps a)-d) on nucleic acids from control samples to identify recurrent sequence variations;   f) applying statistical analysis to determine a confidence interval for the frequency of each sequence variation identified in step e);   g) setting a threshold for the frequency of sequence variations within the confidence interval of step f);   h) eliminating sequence variations whose frequency falls below the threshold set in step g).   
     
     
         18 . A method of assessing a patient by analyzing patient's cell-free nucleic acids by the method of  claim 17 , further comprising step i) of assessing the patient as having cancer if one or more of the sequence variations not eliminated in steps d) and h) are present. 
     
     
         19 . A method of designing a selector comprising a plurality of target genomic regions to be analyzed in a sample of a patient having a type of tumor, the method comprising:
 a) performing sequencing of a genome of the type of tumor from multiple patients;   b) identifying regions of the genome containing a mutation;   c) ranking the regions identified in step b) based on the highest number of patients having a mutation per kilobase of sequence obtained in step a);   d) ranking the regions identified in step b) based on the highest number of patients having a mutation per exon sequenced in step a);   e) including the highest ranked regions from steps c) and d) in the selector.   
     
     
         20 . The method of  claim 19 , wherein the genome sequencing in step a) is exon sequencing. 
     
     
         21 . The method of  claim 19 , wherein regions identified in step b) are at least 100 base pairs long. 
     
     
         22 . The method of  claim 19 , wherein the mutations comprise single nucleotide variations, copy number variations, fusions, seed regions and histology classification regions. 
     
     
         23 . The method of  claim 19 , wherein the highest ranked regions included in the selector comprise the top 10% of the highest ranking regions. 
     
     
         24 . The method of  claim 19 , further comprising after step b), eliminating from the selector regions that fall into repeat-rich regions of the genome. 
     
     
         25 . A method of assessing cancer in a patient comprising:
 a) designing a selector according to  claim 19 ;   b) obtaining a sample from a patient comprising cell-free nucleic acids;   c) determining the sequence of genomic regions of the selector in the patient's nucleic acids;   d) assessing the patient as likely to have cancer or recurrence of cancer if at least one sequence determined in step c) contains a mutation.   
     
     
         26 . The method of  claim 25 , further comprising a confirmation of mutations detected in step b) as somatic in a matched tumor biopsy. 
     
     
         27 . A method of setting a threshold for calling a sequence variant at a base position in a target nucleic acid sequence containing nucleotide substitutions, the method comprising:
 a) performing single molecule sequencing of barcoded nucleic acids from multiple samples to determine the target nucleic acid sequence;   b) for each of the possible classes nucleotide substitutions, determining a total number of substitutions (y) in all positions; a number of supporting reads (t) for the position having a substitution;   c) defining a function relating y to t;   d) solving the function for the desired value of y by determining t, wherein t is the threshold number of reads above which the substitution may be called a variant at the base position in the nucleic acid.   
     
     
         28 . The method of  claim 27 , wherein the threshold t for a given sequence g among the plurality of target sequences is adjusted for global error rate by a method comprising the steps of:
 a) determining error rate e for the plurality of target sequences equal to the number of base positions with nucleotide substitutions in a target sequence divided by the total number of bases in the target sequence;   b) determining sequencing depth d for the plurality of target sequences;   c) if e for sequence g falls within the top 25% of e of the plurality of target sequences, the threshold t for sequence g is adjusted to t′ according to the formula: t′←t×w, where w=min{q 2 , 5} and q=e divided by the 75t h  percentile of the error rates of sequences in the selector;   d) if d for sequence g falls below the median of sequencing depths of the plurality of target sequences (d med ), the threshold t for sequence g is adjusted to t′ according to the formula: t′←t/w*, where w*=ln(d med /d);   
     
     
         29 . A method of assessing a non-small cell lung cancer (NSCLC) patient by analyzing the patient's cfDNA according to  claim 17 , further comprising step i) assessing the patient as assessing the patient as having NSCLC or having a progression of NSCLC if one or more of the sequence variations not eliminated in steps d) and h) are present. 
     
     
         30 . The method of  claim 29 , wherein the mutation is a mutation in epidermal growth factor receptor (EGFR) gene located in the kinase domain (exon 19, 20 and 21) of the gene. 
     
     
         31 . A method of pairing nucleic acid sequencing reads to obtain a double-stranded nucleic acid sequence comprising:
 a) determining the sequence of plurality of single-stranded nucleic comprising insert sequences and adaptor sequences containing barcodes;   b) determining genomic coordinates of the insert sequences;   c) pairing the sequences into a double-stranded nucleic acid if the sequences have complementary barcodes and genomic coordinates of the insert map to the opposite strands.   
     
     
         32 . The method of claim, further comprising a step of eliminating single-member barcode families containing a sequence variant unless the variant is supported by at least one other barcode family with members. 
     
     
         33 . A pool of unique adaptors for analyzing nucleic acids in a sample, each adaptor comprising:
 a double stranded portion at a proximal end and at least one single stranded portion at a distal end, wherein the double stranded portion comprises a double-stranded barcode of at least two base pairs specific to the adaptor, and wherein the single-stranded portion comprises:
 i) a pre-defined single-stranded barcode of at least two nucleotides specific to the sample; and 
 ii) a random single-stranded barcode of at least two nucleotides specific to the adaptor. 
   
     
     
         34 . The pool of unique adaptors of  claim 33 , each adaptor comprising two single-stranded portions at the distal end; one portion comprising a 5′-end and the other portion comprising a 3′-end, wherein the single stranded portions are non-hybridizable with each other. 
     
     
         35 . The pool of unique adaptors of  claim 34 , wherein the two single stranded portions are covalently linked to each other at the distal ends. 
     
     
         36 . The pool of unique adaptors of  claim 35 , wherein the two single stranded portions are covalently linked to each other via a linker. 
     
     
         37 . The pool of unique adaptors of  claim 36 , wherein the linker comprises a cleavage site. 
     
     
         38 . The pool of unique adaptors of  claim 33 , comprising a combination of two sub-pools of adaptors:
 i) a first sub-pool wherein each adaptor comprises two single-stranded portions at the distal end: one portion comprising a 5′-end and the other portion comprising a 3′-end, wherein the single stranded portions are non-hybridizable with each other; and   ii) a second sub-pool wherein each adaptor comprises two non-hybridizable single-stranded portions that are covalently linked to each other at the distal ends.   
     
     
         39 . A method of reduced-error analysis of nucleic acid in a subject's sample comprising:
 a) performing single molecule sequencing nucleic acids from multiple control samples to determine the target nucleic acid sequence;   b) determining the frequency of each of the possible classes of nucleotide substitutions at each position among the control samples;   c) fitting a statistical model to the frequencies determined in step b) to determine frequencies of background errors;   d) performing single molecule sequencing nucleic acids from the subject's sample;   e) determining the frequency of each of the possible classes of nucleotide substitutions at each position in the subject's sample;   f) determining the depth of reads for each target sequence in the subject's sample;   g) applying the statistical model from step c) to the frequencies and depth determined in steps e) and f);   h) eliminating nucleotide substitutions having frequencies below frequencies of background errors determined in step c).

Join the waitlist — get patent alerts

Track US2021363597A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.