US2022349004A1PendingUtilityA1

Compositions and methods for accurately identifying mutations

Assignee: FRED HUTCHINSON CANCER CENTERPriority: Feb 17, 2012Filed: Jun 29, 2022Published: Nov 3, 2022
Est. expiryFeb 17, 2032(~5.6 yrs left)· nominal 20-yr term from priority
Inventors:Jason H. Bielas
C12Q 1/6874C12N 15/1093C12N 15/81C12Q 1/6869C40B 40/08C12N 15/70C40B 50/06C12N 15/1065C12N 15/10C12N 15/85C12Q 1/6827Y02E50/10
83
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.

Claims

exact text as granted — not AI-modified
1 .- 38 . (canceled) 
     
     
         39 . A method of sequencing DNA, the method comprising:
 (a) preparing a sequence library from a sample comprising a plurality of double-stranded DNA molecules from a biological source, wherein preparing the sequence library comprises ligating adapters to the plurality of double-stranded DNA molecules to generate adapter-DNA molecules having a first strand and a second strand;   (b) for each of a plurality of the adapter-DNA molecules:
 (i) sequencing the first and second strands to produce first-strand sequence reads and distinct yet related second-strand sequence reads; and 
 (ii) comparing the first-strand sequence reads and second-strand sequence reads to identify one or more sequence correspondences between the first- and second-strand sequence reads; and 
   (c) comparing the one or more sequence correspondences to a reference sequence to identify a sequence variation as a true mutation, wherein a true mutation is a sequence variation relative to the reference sequence that is consistent between the first-strand sequence reads and second-strand sequence reads for at least one of the adapter-DNA molecules.   
     
     
         40 . The method of  claim 39 , further comprising identifying non-complementary bases between the first-strand sequence reads and the second-strand sequence reads as resulting from experimental or biological error in one strand. 
     
     
         41 . The method of  claim 39 , wherein prior to comparing the first-strand sequence reads and second-strand sequence reads, the method comprises grouping the first-strand sequence reads and second-strand sequence reads based on at least an adapter bar code sequence. 
     
     
         42 . The method of  claim 39 , wherein the adapters comprise bar code sequences selected from a plurality of distinct bar code sequences. 
     
     
         43 . The method of  claim 42 , wherein each bar code sequence is 6 to 8 nucleotides in length. 
     
     
         44 . The method of  claim 42 , wherein each bar code sequence is a double-stranded portion of the respective adapter. 
     
     
         45 . The method of  claim 42 , wherein (i) at least two adapters comprise the same bar code sequence and are ligated to different double-stranded DNA molecules, (ii) end sequences of the double-stranded DNA molecules together with associated bar code sequences form cyphers at both ends of the adapter-DNA molecules; and (iii) individual adapter-target complexes comprise different pairs of cyphers 
     
     
         46 . The method of  claim 39 , wherein the plurality of double-stranded DNA molecules comprises DNA originating from a circulating tumor cell or a cancer cell. 
     
     
         47 . The method of  claim 39 , wherein the sample comprises a tissue obtained from the biological source. 
     
     
         48 . The method of  claim 39 , wherein prior to sequencing, the method further comprises purifying a plurality of adapter-DNA molecules, wherein the purified adapter-DNA molecules comprise nucleic acid molecules that map to specific genomic regions. 
     
     
         49 . The method of  claim 39 , wherein for each of a plurality of the adapter-DNA molecules prior to sequencing, the method further comprises amplifying each original strand of the adapter-DNA molecules to produce a set of copies of the first strand and a set of copies of the second strand. 
     
     
         50 . The method of  claim 39 , further comprising generating an error-corrected sequence, wherein the error-corrected sequence has only nucleotide bases at which the first-strand sequence reads and second-strand sequence reads are in agreement. 
     
     
         51 . A method of sequencing nucleic acid molecules from a biological source, the method comprising:
 (a) providing a sample from the biological source, wherein the sample comprises a plurality of double-stranded nucleic acid molecules,   (b) attaching adapters to individual double-stranded nucleic acid molecules to generate adapter-nucleic acid molecules; and   (c) for each of a plurality of the adapter-nucleic acid molecules:
 (i) generating a set of copies of an original first strand of the adapter-nucleic acid molecule and a set of distinct yet related copies of an original second strand of the adapter-nucleic acid molecule; 
 (ii) sequencing one or more copies of the original first and second strands to produce first- and second-strand sequence reads; 
 (iii) comparing the first- and second-strand sequence reads to identify corresponding base calls that are in agreement; and 
 (iv) generating an error-corrected sequence, wherein the error-corrected sequence has only base calls at which the first- and second-strand sequence reads are in agreement. 
   
     
     
         52 . The method of  claim 51 , further comprising:
 mapping the first- and second-strand sequence reads to a reference sequence to identify sequences corresponding to the reference sequence; and   analyzing sequences corresponding to the reference sequence for each of a plurality of the adapter-nucleic acid molecules to detect the presence of one or more genomic mutations in the sample.   
     
     
         53 . The method of  claim 52 , wherein the one or more genomic mutations comprise a cancer biomarker. 
     
     
         54 . The method of  claim 52 , wherein the biological source is a human, wherein the sample comprises a heterogeneous mixture, and wherein detecting the presence of one or more genomic mutations in the sample characterizes a heterogeneity of the sample. 
     
     
         55 . The method of  claim 54 , wherein the heterogeneous mixture is derived from a tumor. 
     
     
         56 . The method of  claim 54 , wherein the heterogeneous mixture comprises circulating tumor cells or blood. 
     
     
         57 . The method of  claim 51 , wherein the error-corrected sequence comprises a reconstructed sequence of the original first strand of the adapter-nucleic acid molecule and a reconstructed sequence of the original second strand of the adapter-nucleic acid molecule. 
     
     
         58 . The method of  claim 51 , wherein prior to sequencing, the method further comprises purifying a plurality of the adapter-nucleic acid molecules, wherein the purified adapter-nucleic acid molecules comprise nucleic acid molecules that map to specific genomic regions. 
     
     
         59 . The method of  claim 51 , wherein prior to sequencing, the adapter-nucleic acid molecules or copies thereof are selectively enriched by hybridization to substrate bound oligonucleotides. 
     
     
         60 . The method of  claim 51 , wherein prior to comparing the first- and second-strand sequence reads, the method further comprises grouping the first- and second-strand sequence reads based on at least an adapter bar code sequence. 
     
     
         61 . A method of generating sequence reads for a double-stranded target nucleic acid molecule, the method comprising:
 (a) amplifying each original strand of the double-stranded target nucleic acid molecule to produce a set of copies of an original first strand of the double-stranded target nucleic acid molecule and a set of distinct yet related copies of an original second strand of the double-stranded target nucleic acid molecule;   (b) sequencing the amplified target nucleic acid molecules generated from each original strand to produce first- and second-strand sequence reads;   (c) comparing the first- and second-strand sequence reads; and   (d) identifying corresponding base calls in the first- and second-strand sequence reads, wherein corresponding base calls for which first- and second-strand sequence reads are in agreement are identified as true.   
     
     
         62 . The method of  claim 61 , further comprising identifying non-complementary bases between the first- and second-strand sequence reads as experimental errors or sites of DNA damage. 
     
     
         63 . The method of  claim 62 , further comprising generating an error-corrected sequence for the double-stranded target nucleic acid molecule, wherein the error-corrected sequence comprises nucleotide bases at which the first- and second-strand sequence reads are in agreement. 
     
     
         64 . The method of  claim 61 , wherein the double-stranded target nucleic acid molecule comprises a deaminated cytosine. 
     
     
         65 . The method of  claim 64 , wherein the method further comprises enzymatically treating the double-stranded target nucleic acid molecule to repair damaged ends thereof. 
     
     
         66 . The method of  claim 63 , further comprising comparing the error-corrected sequence to a reference sequence to detect the presence of a genetic mutation in the double-stranded target nucleic acid molecule. 
     
     
         67 . The method of  claim 61 , wherein prior to amplifying each original strand of the double-stranded target nucleic acid molecule, the method comprises ligating adapters to each end of the double-stranded target nucleic acid molecule, and wherein the adapters comprise an overhang. 
     
     
         68 . The method of  claim 67 , wherein the adapters comprise bar code sequences, and wherein each bar code sequence is 6 to 8 nucleotides in length.

Join the waitlist — get patent alerts

Track US2022349004A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.