US2022349004A1PendingUtilityA1
Compositions and methods for accurately identifying mutations
Assignee: FRED HUTCHINSON CANCER CENTERPriority: Feb 17, 2012Filed: Jun 29, 2022Published: Nov 3, 2022
Est. expiryFeb 17, 2032(~5.6 yrs left)· nominal 20-yr term from priority
Inventors:Jason H. Bielas
C12Q 1/6874C12N 15/1093C12N 15/81C12Q 1/6869C40B 40/08C12N 15/70C40B 50/06C12N 15/1065C12N 15/10C12N 15/85C12Q 1/6827Y02E50/10
83
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present disclosure provides compositions and methods for accurately detecting mutations by uniquely tagging double stranded nucleic acid molecules with dual cyphers such that sequence data obtained from a sense strand can be linked to sequence data obtained from an anti-sense strand when sequenced, for example, by massively parallel sequencing methods.
Claims
exact text as granted — not AI-modified1 .- 38 . (canceled)
39 . A method of sequencing DNA, the method comprising:
(a) preparing a sequence library from a sample comprising a plurality of double-stranded DNA molecules from a biological source, wherein preparing the sequence library comprises ligating adapters to the plurality of double-stranded DNA molecules to generate adapter-DNA molecules having a first strand and a second strand; (b) for each of a plurality of the adapter-DNA molecules:
(i) sequencing the first and second strands to produce first-strand sequence reads and distinct yet related second-strand sequence reads; and
(ii) comparing the first-strand sequence reads and second-strand sequence reads to identify one or more sequence correspondences between the first- and second-strand sequence reads; and
(c) comparing the one or more sequence correspondences to a reference sequence to identify a sequence variation as a true mutation, wherein a true mutation is a sequence variation relative to the reference sequence that is consistent between the first-strand sequence reads and second-strand sequence reads for at least one of the adapter-DNA molecules.
40 . The method of claim 39 , further comprising identifying non-complementary bases between the first-strand sequence reads and the second-strand sequence reads as resulting from experimental or biological error in one strand.
41 . The method of claim 39 , wherein prior to comparing the first-strand sequence reads and second-strand sequence reads, the method comprises grouping the first-strand sequence reads and second-strand sequence reads based on at least an adapter bar code sequence.
42 . The method of claim 39 , wherein the adapters comprise bar code sequences selected from a plurality of distinct bar code sequences.
43 . The method of claim 42 , wherein each bar code sequence is 6 to 8 nucleotides in length.
44 . The method of claim 42 , wherein each bar code sequence is a double-stranded portion of the respective adapter.
45 . The method of claim 42 , wherein (i) at least two adapters comprise the same bar code sequence and are ligated to different double-stranded DNA molecules, (ii) end sequences of the double-stranded DNA molecules together with associated bar code sequences form cyphers at both ends of the adapter-DNA molecules; and (iii) individual adapter-target complexes comprise different pairs of cyphers
46 . The method of claim 39 , wherein the plurality of double-stranded DNA molecules comprises DNA originating from a circulating tumor cell or a cancer cell.
47 . The method of claim 39 , wherein the sample comprises a tissue obtained from the biological source.
48 . The method of claim 39 , wherein prior to sequencing, the method further comprises purifying a plurality of adapter-DNA molecules, wherein the purified adapter-DNA molecules comprise nucleic acid molecules that map to specific genomic regions.
49 . The method of claim 39 , wherein for each of a plurality of the adapter-DNA molecules prior to sequencing, the method further comprises amplifying each original strand of the adapter-DNA molecules to produce a set of copies of the first strand and a set of copies of the second strand.
50 . The method of claim 39 , further comprising generating an error-corrected sequence, wherein the error-corrected sequence has only nucleotide bases at which the first-strand sequence reads and second-strand sequence reads are in agreement.
51 . A method of sequencing nucleic acid molecules from a biological source, the method comprising:
(a) providing a sample from the biological source, wherein the sample comprises a plurality of double-stranded nucleic acid molecules, (b) attaching adapters to individual double-stranded nucleic acid molecules to generate adapter-nucleic acid molecules; and (c) for each of a plurality of the adapter-nucleic acid molecules:
(i) generating a set of copies of an original first strand of the adapter-nucleic acid molecule and a set of distinct yet related copies of an original second strand of the adapter-nucleic acid molecule;
(ii) sequencing one or more copies of the original first and second strands to produce first- and second-strand sequence reads;
(iii) comparing the first- and second-strand sequence reads to identify corresponding base calls that are in agreement; and
(iv) generating an error-corrected sequence, wherein the error-corrected sequence has only base calls at which the first- and second-strand sequence reads are in agreement.
52 . The method of claim 51 , further comprising:
mapping the first- and second-strand sequence reads to a reference sequence to identify sequences corresponding to the reference sequence; and analyzing sequences corresponding to the reference sequence for each of a plurality of the adapter-nucleic acid molecules to detect the presence of one or more genomic mutations in the sample.
53 . The method of claim 52 , wherein the one or more genomic mutations comprise a cancer biomarker.
54 . The method of claim 52 , wherein the biological source is a human, wherein the sample comprises a heterogeneous mixture, and wherein detecting the presence of one or more genomic mutations in the sample characterizes a heterogeneity of the sample.
55 . The method of claim 54 , wherein the heterogeneous mixture is derived from a tumor.
56 . The method of claim 54 , wherein the heterogeneous mixture comprises circulating tumor cells or blood.
57 . The method of claim 51 , wherein the error-corrected sequence comprises a reconstructed sequence of the original first strand of the adapter-nucleic acid molecule and a reconstructed sequence of the original second strand of the adapter-nucleic acid molecule.
58 . The method of claim 51 , wherein prior to sequencing, the method further comprises purifying a plurality of the adapter-nucleic acid molecules, wherein the purified adapter-nucleic acid molecules comprise nucleic acid molecules that map to specific genomic regions.
59 . The method of claim 51 , wherein prior to sequencing, the adapter-nucleic acid molecules or copies thereof are selectively enriched by hybridization to substrate bound oligonucleotides.
60 . The method of claim 51 , wherein prior to comparing the first- and second-strand sequence reads, the method further comprises grouping the first- and second-strand sequence reads based on at least an adapter bar code sequence.
61 . A method of generating sequence reads for a double-stranded target nucleic acid molecule, the method comprising:
(a) amplifying each original strand of the double-stranded target nucleic acid molecule to produce a set of copies of an original first strand of the double-stranded target nucleic acid molecule and a set of distinct yet related copies of an original second strand of the double-stranded target nucleic acid molecule; (b) sequencing the amplified target nucleic acid molecules generated from each original strand to produce first- and second-strand sequence reads; (c) comparing the first- and second-strand sequence reads; and (d) identifying corresponding base calls in the first- and second-strand sequence reads, wherein corresponding base calls for which first- and second-strand sequence reads are in agreement are identified as true.
62 . The method of claim 61 , further comprising identifying non-complementary bases between the first- and second-strand sequence reads as experimental errors or sites of DNA damage.
63 . The method of claim 62 , further comprising generating an error-corrected sequence for the double-stranded target nucleic acid molecule, wherein the error-corrected sequence comprises nucleotide bases at which the first- and second-strand sequence reads are in agreement.
64 . The method of claim 61 , wherein the double-stranded target nucleic acid molecule comprises a deaminated cytosine.
65 . The method of claim 64 , wherein the method further comprises enzymatically treating the double-stranded target nucleic acid molecule to repair damaged ends thereof.
66 . The method of claim 63 , further comprising comparing the error-corrected sequence to a reference sequence to detect the presence of a genetic mutation in the double-stranded target nucleic acid molecule.
67 . The method of claim 61 , wherein prior to amplifying each original strand of the double-stranded target nucleic acid molecule, the method comprises ligating adapters to each end of the double-stranded target nucleic acid molecule, and wherein the adapters comprise an overhang.
68 . The method of claim 67 , wherein the adapters comprise bar code sequences, and wherein each bar code sequence is 6 to 8 nucleotides in length.Join the waitlist — get patent alerts
Track US2022349004A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.