US2020243161A1PendingUtilityA1

Pathogen detection using next generation sequencing

Assignee: UNIV CALIFORNIAPriority: Sep 21, 2015Filed: Jan 29, 2020Published: Jul 30, 2020
Est. expirySep 21, 2035(~9.1 yrs left)· nominal 20-yr term from priority
G16B 20/20G16B 30/10G16B 30/00G16B 20/00
62
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments are directed to systems and methods for pathogen detection using next-generation sequencing (NGS) analysis of a sample. Embodiments may apply alignment algorithms (e.g., SNAP and/or RAPSearch alignment algorithms) to align individual sequence reads from a sample in a next-generation sequencing (NGS) dataset against reference genome entries in a classified reference genome database. Embodiments of the present invention may include classifying, filtering, and displaying results to a clinician that can then quickly and easily obtain the results of the sequencing to identify a pathogen or other genetic material in a sample that is being tested. A negative sample and a corresponding database can be used to remove contaminants from a list of candidate pathogens. Thus, embodiments are directed to a system that is configured to filter the results of a sequencing alignment and classify a sample quickly.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for identifying one or more pathogens in a sample of biological material, the method comprising, at a computer system:
 receiving a plurality of sequence reads obtained from a sequencing of DNA molecules or reverse transcribed RNA molecules (cDNA) from the sample of biological material, the sample including DNA or cDNA molecules from a plurality of organisms;   using an alignment technique to align the plurality of sequence reads to a plurality of classified reference genomes in a database, thereby obtaining alignment results that include, for each of at least a portion of the sequence reads, at least one matching reference genome to which the sequence read aligns, thereby obtaining matching reference genomes, and wherein the alignment results include classification information for each of the matching reference genomes, the classification information including a plurality of taxonomy identifiers;   identifying a set of the sequence reads that aligned to two or more of the classified reference genomes with at a least a minimum alignment threshold;   for each sequence read of the set of the sequence reads:
 identifying two or more matching reference genomes of the classified reference genomes to which the sequence read aligns with at least the minimum alignment threshold; 
 assigning a taxonomy identifier from the classification information to each of the two or more of the classified reference genomes, thereby obtaining assigned taxonomy identifiers, each taxonomy identifier including at least two levels of classification, the at least two levels having a hierarchy such that there is a lower level and at least one higher level; 
 comparing the assigned taxonomy identifiers of the two or more classified reference genomes at each of the at least two levels of classification; 
 removing each level of the at least two levels from the assigned taxonomy identifiers that do not match between the two or more classified reference genomes; and 
 assigning to the sequence read a lowest level of the at least two levels of the assigned taxonomy identifiers that is shared between the two or more classified reference genomes; and 
 providing an identification of one or more taxonomy identifiers corresponding to one or more candidate pathogens based on numbers of corresponding sequence reads assigned to each of the plurality of taxonomy identifiers. 
   
     
     
         2 . The method of  claim 1 , further comprising:
 when none of the at least two levels of the taxonomy identifiers match, associating an unassigned state with the sequence read.   
     
     
         3 . The method of  claim 1 , wherein providing the identification of the one or more taxonomy identifiers corresponding to the one or more candidate pathogens includes providing an amount of corresponding sequence reads. 
     
     
         4 . The method of  claim 1 , further comprising:
 using sequence reads having an assigned taxonomy identifier to identify the one or more taxonomy identifiers corresponding to the one or more candidate pathogens by:   for each of a set of taxonomy identifiers:
 determining a total score based on at least a coverage of a reference genome corresponding to the taxonomy identifier, thereby obtaining total scores for the set of taxonomy identifiers; 
 ranking the total scores; and 
 identifying the one or more taxonomy identifiers exceeding a score threshold. 
   
     
     
         5 . The method of  claim 4 , wherein the total score is determined further based on:
 a length of the reference genome corresponding to the taxonomy identifier, and   an identity score corresponding to a fractional sum over genomic positions in the reference genome, the fractional sum of a fraction of sequence reads having an exact match at a genomic position relative to all sequence reads covering the genomic position.   
     
     
         6 . The method of  claim 5 , wherein the total score for a taxonomy identifier is determined as (the length+the identity score)*a percent identity score, wherein the percent identity score corresponds to the identity score divided by a number of genomic positions covered by a sequence read aligned to the reference genome. 
     
     
         7 . The method of  claim 1 , wherein the sample of biological material is from a host organism, the method further comprising:
 performing a clinical intervention for the host organism based on the identification of the one or more taxonomy identifiers corresponding to the one or more candidate pathogens.   
     
     
         8 . The method of  claim 1 , further comprising:
 preprocessing the plurality of sequence reads to perform at least one of:   remove sequence reads identified as low quality sequence reads, and   remove sequence reads identified as low complexity sequence reads.   
     
     
         9 . The method of  claim 1 , wherein the classified reference genomes in the database include reference genomes of microorganisms, the method further comprising:
 prior to aligning the plurality of sequence reads to the classified reference genomes in the database:
 aligning the plurality of sequence reads to a reference human genome; and 
 removing a portion of the plurality of sequence reads that align to the reference human genome. 
   
     
     
         10 . The method of  claim 9 , wherein the classified reference genomes in the database include bacterial genomes, fungal genomes, and viral genomes. 
     
     
         11 . The method of  claim 1 , further comprising:
 sequencing DNA molecules from the sample of biological material to obtain at least one million sequence reads, wherein the alignment technique aligns the at least one million sequence reads.   
     
     
         12 . The method of  claim 1 , further comprising filtering the alignment results that are initial alignments results by:
 for each matching reference genome of a set of matching reference genomes:
 based on the initial alignment results, identifying an optimally-aligning sequence read that aligns to the matching reference genome with an optimal alignment score that exceeds or is equal to alignments scores of other sequence reads that align to the matching reference genome, thereby obtaining optimally-aligning sequence reads; 
   for each of the optimally-aligning sequence reads:
 applying a second alignment technique for the optimally-aligning sequence read to the classified reference genomes to obtain a plurality of new alignment scores; 
 determining whether any of the new alignment scores exceeds the optimal alignment score for the optimally-aligning sequence read; 
 when one new alignment score exceeds the optimal alignment score for the optimally-aligning sequence read, removing the matching reference genome from the set of matching reference genomes. 
   
     
     
         13 . A method for identifying pathogens in a test sample of biological material, the method comprising, at a computer system:
 receiving a plurality of sequence reads obtained from a sequencing of DNA molecules from the test sample of biological material, the test sample including DNA molecules from a plurality of organisms;   using an alignment technique to align the plurality of sequence reads to a plurality of classified reference genomes in a reference database, thereby obtaining alignment results that include, for each of at least a portion of the sequence reads, at least one matching reference genome to which the sequence read aligns;   for each matching reference genome of a group of one or more matching reference genomes:
 determining a first amount of sequence reads from the test sample that align to the matching reference genome; 
 determining a second amount of sequence reads from a negative control sample that align to the matching reference genome; 
 determining a ratio of the first amount of sequence reads and the second amount of sequence reads; 
 comparing the ratio to a threshold to determine whether the ratio exceeds the threshold, wherein a set of one or more matching reference genomes have the ratio exceeding the threshold; and 
   providing an output that identifies the set of one or more matching reference genomes as potential pathogens in the test sample.   
     
     
         14 . The method of  claim 13 , further comprising:
 performing the sequencing of DNA or cDNA molecules from the test sample of biological material; and   sequencing DNA or cDNA molecules from the negative control sample, wherein a sequencing of the negative control sample is performed in parallel with the sequencing of the test sample.   
     
     
         15 . The method of  claim 13 , wherein the second amount of sequence reads from the negative control sample is retrieved from a contaminant database. 
     
     
         16 . The method of  claim 15 , wherein an amount of sequence reads for a reference genome in the contaminant database is reduced or removed after a specified time period with the reference genome not being detected in any one of a plurality of negative control samples. 
     
     
         17 . The method of  claim 13 , wherein the threshold is 10. 
     
     
         18 . The method of  claim 13 , further comprising:
 assigning each of the plurality of sequence reads aligned to the at least one matching reference genomes to a single classification level in a taxonomy, wherein the first amount of sequence reads and the second amount of sequence reads are determined for one or more classification levels of the matching reference genome.   
     
     
         19 . The method of  claim 13 , further comprising:
 preparing the negative control sample in parallel with preparing the test sample, the test sample being prepared by purposefully adding DNA or cDNA from a patient sample, wherein preparation of the negative control sample does not include purposefully adding DNA or cDNA;   performing the sequencing of DNA or cDNA molecules from the test sample; and   sequencing any DNA or cDNA molecules from the negative control sample.   
     
     
         20 . The method of  claim 19 , wherein the negative control sample is a buffer used in performing PCR as part of the sequencing of DNA or cDNA molecules from the test sample. 
     
     
         21 . A method for identifying pathogens in a sample of biological material, the method comprising, at a computer system:
 receiving a plurality of sequence reads obtained from a sequencing of DNA molecules from the sample of biological material, the sample including DNA or cDNA molecules from a plurality of organisms;   using an initial alignment technique to align the plurality of sequence reads to a plurality of reference genomes in a database, thereby obtaining initial alignment results that include, for each of at least a portion of the sequence reads, a matching reference genome to which the sequence read aligns;   for each of a set of matching reference genomes:
 based on the initial alignment results, identifying an optimally-aligning sequence read that aligns to the matching reference genome with an optimal alignment score that exceeds or is equal to alignments scores of other sequence reads that align to the matching reference genome, thereby obtaining optimally-aligning sequence reads; 
   for each of the optimally-aligning sequence reads:
 applying a second alignment technique for the optimally-aligning sequence read to the plurality of classified reference genomes to obtain a plurality of new alignment scores; 
   determining whether any of the new alignment scores exceeds the optimal alignment score for the optimally-aligning sequence read;   based on one new alignment score exceeding the optimal alignment score for the optimally-aligning sequence read, removing the matching reference genome from the set of matching reference genomes; and   providing the set of matching reference genomes.   
     
     
         22 . The method of  claim 21 , wherein the matching reference genome is removed from the set of matching reference genomes further based on whether the matching reference genome shares a same taxonomic level as a reference genome corresponding to the one new alignment score. 
     
     
         23 . The method of  claim 21 , further comprising:
 providing the initial alignment results corresponding to the set of matching reference genomes.   
     
     
         24 . The method of  claim 21 , wherein the initial alignment technique uses a global alignment algorithm and the second alignment technique uses a local alignment algorithm. 
     
     
         25 . The method of  claim 24 , wherein the second alignment technique is at least 1,000 times slower than the initial alignment technique.

Join the waitlist — get patent alerts

Track US2020243161A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.