US2019295687A1PendingUtilityA1

Method and system for genome identification

Assignee: COSMOSID INCPriority: Nov 21, 2007Filed: Oct 23, 2018Published: Sep 26, 2019
Est. expiryNov 21, 2027(~1.3 yrs left)· nominal 20-yr term from priority
G16B 50/00G16B 30/00G16B 50/30Y02A90/10
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present invention belongs to the field of genomics and nucleic acid sequencing. It involves a novel method of sequencing biological material and real-time probabilistic matching of short strings of sequencing information to identify all species present in said biological material. It is related to real-time probabilistic matching of sequence information, and more particular to comparing short strings of a plurality of sequences of single molecule nucleic acids, whether amplified or unamplied, whether chemically synthesized or physically interrogated, as fast as the sequence information is generated and in parallel with continuous sequence information generation or collection.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of identifying a biological material in a sample, comprising:
 obtaining a sample comprising said biological material, extracting one or more nucleic acid molecule(s) from said sample, generating sequence information from said nucleic acid molecule(s) with instant direct probabilistic matching for comparison of said sequence information to nucleic acid sequences in a database.   
     
     
         2 . The method of  claim 1 , wherein said one or more nucleic acid molecule(s) is selected from DNA or RNA. 
     
     
         3 . The method of  claim 1 , wherein said sequence information comprises a nucleotide fragment of “n” length. 
     
     
         4 . The method of  claim 3 , wherein said nucleotide fragment of “n” length is compared to the nucleic acid sequences in a database. 
     
     
         5 . The method of  claim 4 , wherein said nucleotide fragment of “n” length is compared to the nucleic acid sequences in a database via probabilistic matching. 
     
     
         6 . The method of  claim 4 , wherein the comparison of said nucleotide fragment of “n” length is performed, in real-time, or as fast as said fragment, or sequence information of said fragment is generated. 
     
     
         7 . The method of  claim 4 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then a nucleic acid fragment of “n+1”, “n+2” . . . “n+x” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein x is less than 50. 
     
     
         8 . The method of  claim 4 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then a nucleic acid fragment of “n+1”, “n+2” . . . “n+x” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is greater than 50. 
     
     
         9 . The method of  claim 1 , further comprising amplification of said one or more nucleic acid molecule(s) to yield a plurality “i” of nucleic acid molecules, prior to generating sequence information. 
     
     
         10 . The method of  claim 8 , wherein said sequence information comprises nucleotide fragments of “n” length. 
     
     
         11 . The method of  claim 9 , wherein the plurality “i” of “n” length nucleotide fragments are compared to the nucleic acid sequences in a database. 
     
     
         12 . The method of  claim 11 , wherein the plurality i(n) of nucleotide fragments are compared to the nucleic acid sequences in a database via probabilistic matching. 
     
     
         13 . The method of  claim 11 , wherein the comparison of plurality i(n) of nucleotide fragments is performed, in real-time, or as fast as said fragments are generated. 
     
     
         14 . The method of  claim 11 , wherein if the probability of match of the plurality i(n) of nucleotide fragments is less than a threshold of a target match, then nucleic acid fragments of “i(n+1)”, “i(n+2)” . . . “i(n+x)” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is less than 50. 
     
     
         15 . The method of  claim 11 , wherein if the probability of match of the plurality i(n) of nucleotide fragments is less than a threshold of a target match, then nucleic acid fragments of “i(n+1)”, “i(n+2)” . . . “i(n+x)” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is greater than 50. 
     
     
         16 . The method according to  claim 5  or  12 , wherein said probabilistic matching is performed using a Bayesian approach. 
     
     
         17 . The method according to  claim 5  or  12 , wherein said probabilistic matching is performed using a Recursive Bayesian approach. 
     
     
         18 . The method according to  claim 5  or  12 , wherein said probabilistic matching is performed using a Naïve Bayesian approach. 
     
     
         19 . The method according to  claim 5  or  12 , wherein said probabilistic matching provides a hierarchical statistical framework to identify the species of said sequence information. 
     
     
         20 . The method of  claim 1 , wherein the comparison of said sequence information to the nucleic acid sequences in a database is performed, in real-time, or as fast as the sequence information is generated, while additional sequence information continues to be generated from said one or more nucleic acid molecule(s). 
     
     
         21 . The method of  claim 20 , wherein said additional sequence information comprises nucleotides of varying lengths. 
     
     
         22 . The method of  claim 1 , wherein said sequence information comprises a nucleotide fragment of “n” length, which is compared, in real-time, or as fast as the fragment is generated to the nucleic acid sequences in a database; while nucleic acid fragments of “n+1”, “n+2” . . . “n+x” length continue to be generated from said one or more nucleic acid molecule(s) and compared, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database. 
     
     
         23 . The method of  claim 1 , wherein said one or more nucleic acid molecule(s) are amplified to yield a plurality “i” of nucleic acid molecules before generating sequence information of “n” length nucleotide fragments; further comprising comparing the plurality i(n) of nucleotide fragments, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database; while a plurality “i(n+1)”, “i(n+2)” . . . “i(n+x)” of nucleic acid fragments continue to be generated from said one or more nucleic acid molecule(s) and compared, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database. 
     
     
         24 . A system for detecting biological material, comprising:
 (i) a sample receiving unit configured to receive a sample comprising biological material;   (ii) an extraction unit in communication with said sample receiving unit, said extraction unit being configured to extract at least one nucleic acid molecule from said sample;   (iii) an sequencing cassette in communication with said extraction unit, said sequencing cassette being configured to receive said at least one nucleic acid molecule from said extraction unit and generate sequence information from said at least one nucleic acid molecule;   (iv) a database comprising reference nucleic acid sequences; and a   (v) processing unit in communication with said sequencing cassette and said database, said processing unit being configured to receive said sequence information from said sequencing cassette and compare said sequence information to said reference nucleic acid sequences.   
     
     
         25 . The system of  claim 24 , comprising:
 a portable sequencing device that electronically transmits data to a database for identification of organisms related to the determination of the sequence of the nucleic acids.   
     
     
         26 . The system of  claim 24 , further comprising a base calling unit configured to processing sequences received by the sequencing cassette. 
     
     
         27 . The system of  claim 26 , wherein the base calling unit is coupled to the probabilistic matching processor. 
     
     
         28 . The system of  claim 27 , wherein the probabilistic matching processor is configured to utilize a Bayesian approach to receive resultant sequence and calculate the probabilities for each sequencing-read while considering sequencing quality scores generated by the base calling unit. 
     
     
         29 . The system of  claim 27 , wherein the probabilistic matching processor uses a database generated and optimized prior to its use for the identification of pathogens. 
     
     
         30 . The system of  claim 27 , wherein the probabilistic matching processor uses weighted scores that vary in accordance to sequence content. 
     
     
         31 . The system of  claim 24 , comprising a storage unit in communication with said processing unit, wherein said processing unit is configured to transmit said sequence information to said data storage unit and subsequently retrieve said sequence information from said data storage unit for processing. 
     
     
         32 . The system of  claim 24 , wherein said at least one nucleic acid molecule is selected from the group consisting of DNA and RNA. 
     
     
         33 . The system of  claim 24 , wherein said sequence information comprises a nucleotide fragment of “n” length. 
     
     
         34 . The system of  claim 33 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences. 
     
     
         35 . The system of  claim 34 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences via probabilistic matching. 
     
     
         36 . The system of  claim 34 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences in real time, or as fast as said fragment of “n” length is generated. 
     
     
         37 . The system of  claim 34 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then said sequencing cassette is configured to generate sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length from said one or more nucleic acid molecule(s) and said extraction unit is configured to compared said nucleotide fragments of “n+1”, “n+2” . . . “n+x” length to the nucleic acid sequences in a database. 
     
     
         38 . The system of  claim 36 , wherein said nucleotide fragment of “n” length is compared to said reference nucleic acid sequences in real time, or as fast as said fragment of “n” length is generated, while the sequencing unit continues to generate sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length from said one or more nucleic acid molecule(s), and the processing unit compares said sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database. 
     
     
         39 . A method of identifying a biological material in a sample, comprising:
 (i) obtaining a sample comprising said biological material,   (ii) extracting one or more nucleic acid molecule(s) from said sample,   (iii) generating sequence information, comprising a sequence of a nucleotide fragment from said one or more nucleic acid molecule(s),   (iv) comparing said sequence of a nucleotide fragment to nucleic acid sequences in a database;   
       and if said comparison of said sequence of a nucleotide fragment does not result in a match identifying the biological material in said sample, then the method further comprises:
 (v) generating additional sequence information from said one or more nucleic acid molecule(s), wherein said additional sequence information comprises a sequence of a nucleotide fragment consisting of one additional nucleotide, 
 (vi) comparing said additional sequence information to nucleic acid sequences in a database immediately following the generation of said additional sequence information, 
 
       and repeating steps (v)-(vi) until a match results in the identification of the biological material is said sample. 
     
     
         40 . A method of identifying a biological material in a sample, comprising:
 (i) obtaining a sample comprising said biological material,   (ii) extracting one or more nucleic acid molecule(s) from said sample,   (iii) amplifying said one or more nucleic acid molecule(s) to yield a plurality of one or more nucleic acid molecule(s),   (iii) generating a plurality of sequence information, comprising a plurality of sequences of a nucleotide fragment, from said plurality of one or more nucleic acid molecule(s),   (iv) comparing said plurality of sequences of a nucleotide fragment to nucleic acid sequences in a database,   
       and if said comparison of said plurality of sequences of a nucleotide fragment does not result in a match identifying the biological material in said sample, then the method further comprises:
 (v) generating plurality of additional sequence information from said one or more nucleic acid molecule(s), wherein said additional sequence information comprises a sequence of a nucleotide fragment consisting of one additional nucleotide, 
 (vi) comparing said additional sequence information to nucleic acid sequences in a database immediately following the generation of said additional sequence information, 
 
       and repeating steps (v)-(vi) until a match results in the identification of the biological material is said sample. 
     
     
         41 . The methods of  claim 39  or  40 , wherein the comparison to the nucleic acid sequences in a database is performed via probabilistic matching as fast as the sequence information is generated.

Join the waitlist — get patent alerts

Track US2019295687A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.