US2022348907A1PendingUtilityA1

Methods for enriching for duplex reads in sequencing and error correction

Assignee: GRAIL INCPriority: Dec 15, 2017Filed: Jun 29, 2022Published: Nov 3, 2022
Est. expiryDec 15, 2037(~11.4 yrs left)· nominal 20-yr term from priority
C40B 80/00C12Q 1/6855C40B 50/10C12Y 207/07007C40B 40/06C12Q 1/6806C12Q 1/6848C12N 15/1093C40B 50/04C12N 15/1072G16B 30/10C12N 15/1089C12N 15/66G16B 35/00
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods for preparing sequencing libraries from a DNA-containing test sample, as well as methods for correcting sequencing-derived errors, are provided.

Claims

exact text as granted — not AI-modified
1 .- 29 . (canceled) 
     
     
         30 . A method for preparing a sequencing library from a test sample comprising a plurality of double-stranded DNA molecules, the method comprising:
 (a) obtaining a test sample comprising a plurality of double-stranded DNA (dsDNA) molecules, wherein the dsDNA molecules comprise a forward strand sequence and a reverse complement strand sequence;   (b) ligating dsDNA adapters to both ends of the dsDNA molecules, generating a plurality of dsDNA adapter-molecule constructs, wherein the dsDNA adapter comprises a unique molecular identifier (UMI);   (c) incorporating one or more biotin-labeled nucleotides into partially ligated or nicked dsDNA adapter-molecule constructs to create a plurality of labeled dsDNA adapter-molecule constructs;   (d) depleting the labeled dsDNA adapter-molecule constructs from the test sample;   (e) amplifying the remaining dsDNA adapter-molecule constructs in the depleted test sample to generate a sequencing library;   (f) sequencing at least a portion of the sequencing library to obtain a plurality of sequence reads;   (g) grouping the sequence reads into families based on the UMIs, wherein each family comprises a first set of forward strand sequences each having a first UMI and a second set of reverse complement strand sequences each having a second UMI; and   (h) comparing the sequence reads within each family to generate a consensus sequence for each family, thereby correcting sequencing-derived errors in sequence reads.   
     
     
         31 . The method according to  claim 30 , further comprising:
 (i) aligning the consensus sequences to a reference sequence and identifying consensus sequences as one or more rare variants if the one or more consensus sequences vary from the reference sequence at one or more nucleotide positions.   
     
     
         32 . The method according to  claim 30 , wherein the dsDNA molecules are cell-free DNA (cfDNA) molecules. 
     
     
         33 . The method according to  claim 32 , wherein the cfDNA molecules originate from healthy cells and from cancer cells. 
     
     
         34 . The method according to  claim 30 , wherein the test sample comprises whole blood, a blood fraction, plasma, serum, urine, fecal matter, saliva, a tissue biopsy, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or any combination thereof. 
     
     
         35 . The method according to  claim 30 , wherein the plurality of dsDNA molecules are modified prior to adapter ligation, and wherein the modification comprises end-repairing and A-tailing prior to adapter ligation. 
     
     
         36 . The method according to of  claim 30 , wherein the adapters further comprise a sample-specific index sequence. 
     
     
         37 . The method according to  claim 30 , wherein the adapters further comprise a universal priming site. 
     
     
         38 . The method according to  claim 30 , wherein the adapters further comprise one or more sequencing oligonucleotides for use in cluster generation and/or sequencing. 
     
     
         39 . The method according to  claim 30 , wherein one or more biotin-labeled nucleotides are incorporated into partially ligated or nicked dsDNA adapter-molecule constructs using a DNA polymerase. 
     
     
         40 . The method according to  claim 39 , wherein the DNA polymerase is a DNA polymerase comprising strand displacement activity. 
     
     
         41 . The method according to  claim 39 , wherein the DNA polymerase lacks exonuclease activity. 
     
     
         42 . The method according to  claim 39 , wherein the DNA polymerase is Bacillus stearothermophilus DNA polymerase (Bst Pol), a Klenow DNA polymerase, or a phi29 DNA polymerase. 
     
     
         43 . The method according to  claim 42 , wherein the DNA polymerase is a Klenow DNA polymerase that lacks exonuclease activity. 
     
     
         44 . The method according to  claim 30 , wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each base is identified at a given position in the sequence when a specific base is present in a majority of the sequence reads of the family. 
     
     
         45 . The method according to  claim 30 , wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each base is identified at a given position in the sequence when a specific base is present in at least 70% of the sequence reads comprising the family. 
     
     
         46 . The method according to  claim 30 , wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each base is identified at a given position in the sequence when a specific base is present in at least 80% of the sequence reads comprising the family. 
     
     
         47 . The method according to  claim 30 , wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each base is identified at a given position in the sequence when a specific base is present in at least 90% of the sequence reads comprising the family. 
     
     
         48 . The method according to  claim 30 , wherein the consensus sequence comprises a sequence of nucleotide bases, wherein each base is identified at a given position in the sequence when a specific base is present in at least 95% of the sequence reads comprising the family. 
     
     
         49 . The method according to  claim 30 , wherein the method further comprises loading at least a portion of the sequencing library into a sequencing flow cell and generating a plurality of sequencing clusters on the flow cell, wherein the clusters comprise the forward strand sequence and the reverse complement strand sequence. 
     
     
         50 . The method according to  claim 30 , wherein the sequence reads are obtained from next-generation sequencing (NGS), massively parallel sequencing using sequencing-by-synthesis, paired-end sequencing, or any combination thereof. 
     
     
         51 . The method according to  claim 50 , wherein the sequence reads are obtained from paired-end sequencing, where the sequence reads comprise a read pair, and wherein each read pair comprises a first read of the forward strand sequence and second read of the reverse complement strand sequence. 
     
     
         52 . The method according to  claim 31 , wherein the method further comprises using the one or more rare variants to detect the presence or absence of a cancer, determine a cancer status, monitor cancer progression, and/or determine a cancer classification. 
     
     
         53 . The method according to  claim 52 , wherein monitoring cancer progression comprises monitoring disease progression, monitoring therapy, or monitoring cancer growth. 
     
     
         54 . The method according to  claim 52 , wherein the cancer classification comprises determining a cancer type and/or a cancer tissue of origin. 
     
     
         55 . The method according to  claim 52 , wherein the cancer comprises a carcinoma, a sarcoma, a myeloma, a leukemia, a lymphoma, a blastoma, a germ cell tumor, or any combination thereof.

Join the waitlist — get patent alerts

Track US2022348907A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.