US2021358567A1PendingUtilityA1

Systems and methods for detecting insertions and deletions

Assignee: GUARDANT HEALTH INCPriority: Sep 30, 2016Filed: Jul 22, 2021Published: Nov 18, 2021
Est. expirySep 30, 2036(~10.2 yrs left)· nominal 20-yr term from priority
Inventors:Marcin Sikora
G16B 30/00G16B 20/20C12Q 2535/122C12Q 2537/159C12Q 2600/156C12Q 2600/154G16B 20/50C12Q 1/6886C12Q 1/6806C12Q 1/6827G16B 25/10G16B 5/00G16B 40/30
79
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods are provided for improving accuracy of detecting an insertion or deletion (indel) from a plurality of sequence reads derived from cell-free deoxyribonucleic acid (cfDNA). For each of the plurality of sequence reads associated with cfDNA molecules, a candidate indel may be identified. Each candidate indel may then be classified as either a true indel or an introduced indel, using a combination of predetermined expectations of (i) an indel being detected in one or more sequence reads, (ii) that a detected indel is a true indel present in a given cfDNA molecule, given that an indel has been detected in the one or more of the sequence reads, and/or (iii) that a detected indel is introduced by non-biological error, given that an indel has been detected in the one or more of the sequence reads, in conjunction with one or more model parameters to perform a hypothesis test.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory computer-readable medium comprising machine executable code that, upon execution by one or more computer processors, implements a method for improving accuracy of detecting an insertion or deletion (indel) from a plurality of sequence reads derived from cell-free deoxyribonucleic acid (cfDNA) molecules in a bodily sample of a subject, which plurality of sequence reads are generated by nucleic acid sequencing, comprising:
 (a) for each of the plurality of sequence reads associated with the cfDNA molecules, providing one or more of:
 a predetermined expectation of an indel being detected in one or more sequence reads of the plurality of sequence reads; 
 a predetermined expectation that a detected indel is a true indel present in a given cfDNA molecule of the cfDNA molecules, given that an indel has been detected in the one or more of the sequence reads; and 
 a predetermined expectation that a detected indel is introduced by non-biological error, given that an indel has been detected in the one or more of the sequence reads; 
   (b) providing quantitative measures of one or more model parameters characteristic of sequence reads generated by nucleic acid sequencing;   (c) detecting one or more candidate indels in the plurality of sequence reads associated with the cfDNA molecules; and   (d) for each candidate indel, performing a hypothesis test using one or more of the model parameters in conjunction with one or more of the predetermined expectations to classify said candidate indel as a true indel or an introduced indel, thereby improving accuracy of detecting an indel.   
     
     
         2 . The non-transitory computer-readable medium of  claim 1 , comprising providing a predetermined expectation of an indel being detected in one or more sequence reads of the plurality of sequence reads. 
     
     
         3 . The non-transitory computer-readable medium of  claim 1 , comprising providing:
 a predetermined expectation of an indel being detected in one or more sequence reads of the plurality of sequence reads;   a predetermined expectation that a detected indel is a true indel present in a given cfDNA molecule of the cfDNA molecules, given that an indel has been detected in the one or more of the sequence reads; and   a predetermined expectation that a detected indel is introduced by non-biological error, given that an indel has been detected in the one or more of the sequence reads.   
     
     
         4 . The non-transitory computer-readable medium of  claim 3 , wherein for each candidate indel, performing a hypothesis test using one or more of the model parameters in conjunction with each of the predetermined expectations to classify said candidate indel as a true indel or an introduced indel, thereby improving accuracy of detecting an indel. 
     
     
         5 . The non-transitory computer-readable medium of  claim 1 , wherein the non-biological error comprises error in sequencing at a plurality of genomic base locations. 
     
     
         6 . The non-transitory computer-readable medium of  claim 1 , wherein the non-biological error comprises error in amplification at a plurality of genomic base locations. 
     
     
         7 . The non-transitory computer-readable medium of  claim 1 , wherein the model parameters comprise for each of one or more variant alleles, a frequency of the variant allele (α) and a frequency of non-reference alleles other than the variant allele (α′). 
     
     
         8 . The non-transitory computer-readable medium of  claim 1 , wherein the model parameters comprise one or more of:
 (i) for each of one or more variant alleles, a frequency of the variant allele (α) and a frequency of non-reference alleles other than the variant allele (α′).   (ii) a frequency of an indel error in the entire forward strand of a family of strands (β 1 ), wherein a family comprises a collection of amplicons originating from a single strand of the cell free DNA molecules;   (iii) a frequency of an indel error in the entire reverse strand of a family of strands (β 2 ); and   (iv) a frequency of an indel error in a sequence read (γ).   
     
     
         9 . The non-transitory computer-readable medium of  claim 1 , wherein the step of performing a hypothesis test comprises performing a multi-parameter maximization algorithm. 
     
     
         10 . The non-transitory computer-readable medium of  claim 9 , wherein the multi-parameter maximization algorithm comprises a Nelder-Mead algorithm. 
     
     
         11 . The non-transitory computer-readable medium of  claim 1 , wherein the classifying of (d) comprises:
 (i) maximizing a multi-parameter likelihood function,   (ii) classifying a candidate indel as a true indel if the maximum likelihood function value is greater than a predetermined threshold value, and   (iii) classifying a candidate indel as an introduced indel if the maximum likelihood function value is less than or equal to a predetermined threshold value.   
     
     
         12 . The non-transitory computer-readable medium of  claim 1 , wherein the sequence reads are enriched for one or more loci. 
     
     
         13 . The non-transitory computer-readable medium of  claim 1 , wherein the cfDNA molecules comprise circulating tumor DNA. 
     
     
         14 . The non-transitory computer-readable medium of  claim 1 , wherein the bodily sample comprises whole blood, platelets, serum, plasma, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine, cervical fluid or lavage, vaginal fluid or lavage, mammary gland or lavage. 
     
     
         15 . The non-transitory computer-readable medium of  claim 1 , wherein the bodily sample comprises plasma. 
     
     
         16 . The non-transitory computer-readable medium of  claim 1 , wherein the subject is an individual that has or is suspected of having a disease or a pre-disposition to the disease. 
     
     
         17 . The non-transitory computer-readable medium of  claim 1 , wherein the subject is an individual that is in need of therapy or suspected of needing therapy. 
     
     
         18 . The non-transitory computer-readable medium of  claim 1 , wherein detecting one or more candidate indels in the plurality of sequence reads associated with the cfDNA molecules comprises aligning to a reference sequence. 
     
     
         19 . A system for improving accuracy of detecting an insertion or deletion (indel) from a plurality of sequence reads derived from cell-free deoxyribonucleic acid (cfDNA) molecules in a bodily sample of a subject, which plurality of sequence reads are generated by nucleic acid sequencing, the system comprising:
 (a) one or more computer processors; and   (b) a non-transitory computer-readable medium comprising machine-executable code that, upon execution by the one or more computer processors, implements a method comprising:
 (i) for each of the plurality of sequence reads associated with the cfDNA molecules, providing one or more of: 
 a predetermined expectation of an indel being detected in one or more sequence reads of the plurality of sequence reads; 
 a predetermined expectation that a detected indel is a true indel present in a given cfDNA molecule of the cfDNA molecules, given that an indel has been detected in the one or more of the sequence reads; and 
 a predetermined expectation that a detected indel is introduced by non-biological error, given that an indel has been detected in the one or more of the sequence reads; 
 (ii) providing quantitative measures of one or more model parameters characteristic of sequence reads generated by nucleic acid sequencing; 
 (iii) detecting one or more candidate indels in the plurality of sequence reads associated with the cfDNA molecules; and 
 (iv) for each candidate indel, performing a hypothesis test using one or more of the model parameters in conjunction with one or more of the predetermined expectations to classify said candidate indel as a true indel or an introduced indel, thereby improving accuracy of detecting an indel. 
   
     
     
         20 . The system of  claim 19 , further comprising a communication interface that receives, over a communication network, the sequence reads generated by the nucleic acid sequencing.

Join the waitlist — get patent alerts

Track US2021358567A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.