US2021317525A1PendingUtilityA1

Compositions and methods for accurately identifying mutations

Assignee: HUTCHINSON FRED CANCER RESPriority: Feb 17, 2012Filed: Jun 23, 2021Published: Oct 14, 2021
Est. expiryFeb 17, 2032(~5.6 yrs left)· nominal 20-yr term from priority
Inventors:Jason H. Bielas
C12N 15/10C12N 15/81C40B 50/06C12Q 1/6869C12N 15/1093C12Q 1/6874C12N 15/70C12N 15/1065C40B 40/08C12N 15/85C12Q 1/6827Y02E50/10
79
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.

Claims

exact text as granted — not AI-modified
1 .- 38 . (canceled) 
     
     
         39 . A method comprising:
 (a) providing a plurality of circulating DNA molecules obtained from a patient sample;   (b) ligating the circulating DNA molecules to cypher polynucleotides to form double-stranded cypher-target nucleic acid complexes, wherein:
 (i) the cypher polynucleotides comprise identifier tags selected from a plurality of distinct identifier tag sequences; 
 (ii) at least two of the identifier tags are identical in sequence and are ligated to different circulating DNA molecules, thereby non-uniquely tagging the different circulating DNA molecules; and 
 (iii) an identifier tag alone or in combination with an end of a circulating DNA molecule uniquely identifies a cypher-target nucleic acid complex; 
   (c) amplifying the cypher-target nucleic acid complexes to produce a corresponding plurality of cypher-target amplification products;   (d) sequencing the cypher-target amplification products to produce a plurality of sequencing reads;   (e) grouping the sequencing reads into groups, each of the groups comprising the same identifier tag sequence and the same circulating DNA end sequences, wherein each of the groups comprises sequencing reads from the cypher-target amplification products of one of the cypher-target nucleic acid complexes; and   (f) comparing the sequencing reads within the groups, and generating error-corrected sequences of the circulating DNA molecules by distinguishing erroneous nucleotides in one strand that lack a matched base change in the complementary strand.   
     
     
         40 . The method of  claim 39 , further comprising identifying one or more single nucleotide mutations. 
     
     
         41 . The method of  claim 39 , wherein the ligating comprises ligating to an overhang or a blunt end. 
     
     
         42 . The method of  claim 39 , further comprising detecting mutations in one or more of the error-corrected sequences as compared to a reference sequence. 
     
     
         43 . The method of  claim 39 , wherein sequencing the cypher-target amplification products comprises converting data from a sequencing instrument into quality scores and then into sequencing reads. 
     
     
         44 . The method of  claim 39 , further comprising purifying a plurality of cypher-target nucleic acid complexes prior to sequencing, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules from specific genomic regions. 
     
     
         45 . The method of  claim 39 , wherein the plurality of circulating DNA molecules comprise a mutation present at a frequency of 2.1×10 −6  or lower. 
     
     
         46 . The method of  claim 39 , wherein generating the error corrected sequences results in a measureable sequencing error rate from about 10 −6  to about 10 −8 . 
     
     
         47 . The method of  claim 39 , wherein the circulating DNA molecules comprise plasma DNA biomarkers. 
     
     
         48 . The method of  claim 39 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of about 5 nucleotides in length. 
     
     
         49 . The method of  claim 39 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of 5 or 6 nucleotides in length. 
     
     
         50 . The method of  claim 39 , wherein the cypher polynucleotides comprising the identifier tags are contained within a pool of cypher polynucleotides comprising known sequences. 
     
     
         51 . The method of  claim 39 , wherein the ligating comprises ligating identifier tags to both ends of the circulating DNA molecules, and further wherein the identifier tags at both ends together form a unique pair of identifiers that differ between each of the other pairs of identifiers ligated to the circulating DNA molecules. 
     
     
         52 . The method of  claim 39 , wherein grouping sequencing reads is based on (i) the identifier tag sequence and (ii) sequence information from an end of the circulating DNA molecule. 
     
     
         53 . The method of  claim 39 , wherein:
 (i) the plurality of cypher-target amplification products comprises amplification products from first strands and complementary second strands of the cypher-target nucleic acid complexes;   (ii) the plurality of sequencing reads comprises a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads; and   (iii) the comparing comprises comparing the first-strand sequencing reads with the second-strand sequencing reads within the groups.   
     
     
         54 . A method comprising:
 (a) ligating cypher polynucleotides to circulating DNA molecules obtained from a patient sample to form double-stranded cypher-target nucleic acid complexes, wherein:
 (i) the cypher polynucleotides comprise identifier tags selected from a plurality of distinct identifier tag sequences; 
 (ii) at least two of the identifier tags are identical in sequence and are ligated to different circulating DNA molecules, thereby non-uniquely tagging the different circulating DNA molecules; and 
 (iii) an identifier tag alone or in combination with an end of a circulating DNA molecule uniquely identifies a cypher-target nucleic acid complex; 
   (b) amplifying the cypher-target nucleic acid complexes to produce a corresponding plurality of cypher-target amplification products;   (c) sequencing the cypher-target amplification products to produce a plurality of sequencing reads;   (d) grouping the sequencing reads based on (i) the identifier tag sequence and (ii) sequence information from the circulating DNA molecule, wherein a group comprises sequencing reads from the cypher-target amplification products of one of the cypher-target nucleic acid complexes; and   (e) comparing the sequencing reads within the groups, and generating error-corrected sequences of the circulating DNA molecules by distinguishing erroneous nucleotides in one strand that lack a matched base change in the complementary strand.   
     
     
         55 . The method of  claim 54 , further comprising purifying a plurality of cypher-target nucleic acid complexes prior to sequencing, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules from specific genomic regions. 
     
     
         56 . The method of  claim 54 , further comprising identifying one or more single nucleotide mutations. 
     
     
         57 . The method of  claim 54 , wherein the circulating DNA molecules comprise blood biomarkers. 
     
     
         58 . The method of  claim 54 , wherein the circulating DNA molecules comprise DNA molecules derived from cancer cells. 
     
     
         59 . The method of  claim 54 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of about 5 nucleotides in length. 
     
     
         60 . The method of  claim 54 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of 5 or 6 nucleotides in length. 
     
     
         61 . The method of  claim 54 , wherein:
 (i) the plurality of cypher-target amplification products comprises amplification products from first strands and complementary second strands of the cypher-target nucleic acid complexes;   (ii) the plurality of sequencing reads comprises a plurality of first-strand sequencing reads and a plurality of second-strand sequencing reads; and   (iii) the comparing comprises comparing the first-strand sequencing reads with the second-strand sequencing reads within the groups.   
     
     
         62 . The method of  claim 53 , further comprising identifying one or more single nucleotide mutations. 
     
     
         63 . The method of  claim 53 , wherein the ligating comprises ligating to an overhang or a blunt end. 
     
     
         64 . The method of  claim 53 , further comprising detecting mutations in one or more of the error-corrected sequences as compared to a reference sequence. 
     
     
         65 . The method of  claim 53 , wherein sequencing the cypher-target amplification products comprises converting data from a sequencing instrument into quality scores and then into sequencing reads. 
     
     
         66 . The method of  claim 53 , further comprising purifying a plurality of cypher-target nucleic acid complexes prior to sequencing, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules from specific genomic regions. 
     
     
         67 . The method of  claim 53 , wherein the plurality of circulating DNA molecules comprise a mutation present at a frequency of 2.1×10 −6  or lower. 
     
     
         68 . The method of  claim 53 , wherein generating the error corrected sequences results in a measureable sequencing error rate from about 10 −6  to about 10 −8 . 
     
     
         69 . The method of  claim 53 , wherein the circulating DNA molecules comprise plasma DNA biomarkers. 
     
     
         70 . The method of  claim 53 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of about 5 nucleotides in length. 
     
     
         71 . The method of  claim 53 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of 5 or 6 nucleotides in length. 
     
     
         72 . The method of  claim 53 , wherein the cypher polynucleotides comprising the identifier tags are contained within a pool of cypher polynucleotides comprising known sequences. 
     
     
         73 . The method of  claim 53 , wherein the ligating comprises ligating identifier tags to both ends of the circulating DNA molecules, and further wherein the identifier tags at both ends together form a unique pair of identifiers that differ between each of the other pairs of identifiers ligated to the circulating DNA molecules. 
     
     
         74 . The method of  claim 53 , wherein grouping sequencing reads is based on (i) the identifier tag sequence and (ii) sequence information from an end of the circulating DNA molecule. 
     
     
         75 . The method of  claim 61 , further comprising purifying a plurality of cypher-target nucleic acid complexes prior to sequencing, wherein the purified cypher-target nucleic acid complexes comprise nucleic acid molecules from specific genomic regions. 
     
     
         76 . The method of  claim 61 , further comprising identifying one or more single nucleotide mutations. 
     
     
         77 . The method of  claim 61 , wherein the circulating DNA molecules comprise blood biomarkers. 
     
     
         78 . The method of  claim 61 , wherein the circulating DNA molecules comprise DNA molecules derived from cancer cells. 
     
     
         79 . The method of  claim 61 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of about 5 nucleotides in length. 
     
     
         80 . The method of  claim 61 , wherein each identifier tag of the plurality of distinct identifier tag sequences is a random or partially random sequence of 5 or 6 nucleotides in length.

Join the waitlist — get patent alerts

Track US2021317525A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.