Method and system for genome identification
Abstract
The present invention belongs to the field of genomics and nucleic acid sequencing. It involves a novel method of sequencing biological material and real-time probabilistic matching of short strings of sequencing information to identify all species present in said biological material. It is related to real-time probabilistic matching of sequence information, and more particular to comparing short strings of a plurality of sequences of single molecule nucleic acids, whether amplified or unamplied, whether chemically synthesized or physically interrogated, as fast as the sequence information is generated and in parallel with continuous sequence information generation or collection.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of identifying a biological material in a sample, comprising:
obtaining a sample comprising said biological material, extracting one or more nucleic acid molecule(s) from said sample, generating sequence information from said nucleic acid molecule(s) with instant direct probabilistic matching for comparison of said sequence information to nucleic acid sequences in a database.
2 . The method of claim 1 , wherein said one or more nucleic acid molecule(s) is selected from DNA or RNA.
3 . The method of claim 1 , wherein said sequence information comprises a nucleotide fragment of “n” length.
4 . The method of claim 3 , wherein said nucleotide fragment of “n” length is compared to the nucleic acid sequences in a database.
5 . The method of claim 4 , wherein said nucleotide fragment of “n” length is compared to the nucleic acid sequences in a database via probabilistic matching.
6 . The method of claim 4 , wherein the comparison of said nucleotide fragment of “n” length is performed, in real-time, or as fast as said fragment, or sequence information of said fragment is generated.
7 . The method of claim 4 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then a nucleic acid fragment of “n+1”, “n+2” . . . “n+x” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein x is less than 50.
8 . The method of claim 4 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then a nucleic acid fragment of “n+1”, “n+2” . . . “n+x” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is greater than 50.
9 . The method of claim 1 , further comprising amplification of said one or more nucleic acid molecule(s) to yield a plurality “i” of nucleic acid molecules, prior to generating sequence information.
10 . The method of claim 8 , wherein said sequence information comprises nucleotide fragments of “n” length.
11 . The method of claim 9 , wherein the plurality “i” of “n” length nucleotide fragments are compared to the nucleic acid sequences in a database.
12 . The method of claim 11 , wherein the plurality i(n) of nucleotide fragments are compared to the nucleic acid sequences in a database via probabilistic matching.
13 . The method of claim 11 , wherein the comparison of plurality i(n) of nucleotide fragments is performed, in real-time, or as fast as said fragments are generated.
14 . The method of claim 11 , wherein if the probability of match of the plurality i(n) of nucleotide fragments is less than a threshold of a target match, then nucleic acid fragments of “i(n+1)”, “i(n+2)” . . . “i(n+x)” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is less than 50.
15 . The method of claim 11 , wherein if the probability of match of the plurality i(n) of nucleotide fragments is less than a threshold of a target match, then nucleic acid fragments of “i(n+1)”, “i(n+2)” . . . “i(n+x)” length is generated from said one or more nucleic acid molecule(s) and compared to the nucleic acid sequences in a database, wherein “x” is greater than 50.
16 . The method according to claim 5 or 12 , wherein said probabilistic matching is performed using a Bayesian approach.
17 . The method according to claim 5 or 12 , wherein said probabilistic matching is performed using a Recursive Bayesian approach.
18 . The method according to claim 5 or 12 , wherein said probabilistic matching is performed using a Naïve Bayesian approach.
19 . The method according to claim 5 or 12 , wherein said probabilistic matching provides a hierarchical statistical framework to identify the species of said sequence information.
20 . The method of claim 1 , wherein the comparison of said sequence information to the nucleic acid sequences in a database is performed, in real-time, or as fast as the sequence information is generated, while additional sequence information continues to be generated from said one or more nucleic acid molecule(s).
21 . The method of claim 20 , wherein said additional sequence information comprises nucleotides of varying lengths.
22 . The method of claim 1 , wherein said sequence information comprises a nucleotide fragment of “n” length, which is compared, in real-time, or as fast as the fragment is generated to the nucleic acid sequences in a database; while nucleic acid fragments of “n+1”, “n+2” . . . “n+x” length continue to be generated from said one or more nucleic acid molecule(s) and compared, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database.
23 . The method of claim 1 , wherein said one or more nucleic acid molecule(s) are amplified to yield a plurality “i” of nucleic acid molecules before generating sequence information of “n” length nucleotide fragments; further comprising comparing the plurality i(n) of nucleotide fragments, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database; while a plurality “i(n+1)”, “i(n+2)” . . . “i(n+x)” of nucleic acid fragments continue to be generated from said one or more nucleic acid molecule(s) and compared, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database.
24 . A system for detecting biological material, comprising:
(i) a sample receiving unit configured to receive a sample comprising biological material; (ii) an extraction unit in communication with said sample receiving unit, said extraction unit being configured to extract at least one nucleic acid molecule from said sample; (iii) an sequencing cassette in communication with said extraction unit, said sequencing cassette being configured to receive said at least one nucleic acid molecule from said extraction unit and generate sequence information from said at least one nucleic acid molecule; (iv) a database comprising reference nucleic acid sequences; and a (v) processing unit in communication with said sequencing cassette and said database, said processing unit being configured to receive said sequence information from said sequencing cassette and compare said sequence information to said reference nucleic acid sequences.
25 . The system of claim 24 , comprising:
a portable sequencing device that electronically transmits data to a database for identification of organisms related to the determination of the sequence of the nucleic acids.
26 . The system of claim 24 , further comprising a base calling unit configured to processing sequences received by the sequencing cassette.
27 . The system of claim 26 , wherein the base calling unit is coupled to the probabilistic matching processor.
28 . The system of claim 27 , wherein the probabilistic matching processor is configured to utilize a Bayesian approach to receive resultant sequence and calculate the probabilities for each sequencing-read while considering sequencing quality scores generated by the base calling unit.
29 . The system of claim 27 , wherein the probabilistic matching processor uses a database generated and optimized prior to its use for the identification of pathogens.
30 . The system of claim 27 , wherein the probabilistic matching processor uses weighted scores that vary in accordance to sequence content.
31 . The system of claim 24 , comprising a storage unit in communication with said processing unit, wherein said processing unit is configured to transmit said sequence information to said data storage unit and subsequently retrieve said sequence information from said data storage unit for processing.
32 . The system of claim 24 , wherein said at least one nucleic acid molecule is selected from the group consisting of DNA and RNA.
33 . The system of claim 24 , wherein said sequence information comprises a nucleotide fragment of “n” length.
34 . The system of claim 33 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences.
35 . The system of claim 34 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences via probabilistic matching.
36 . The system of claim 34 , wherein said extraction unit is configured to compare said nucleotide fragment of “n” length to said reference nucleic acid sequences in real time, or as fast as said fragment of “n” length is generated.
37 . The system of claim 34 , wherein if the probability of match of a nucleotide fragment of “n” length is less than a threshold of a target match, then said sequencing cassette is configured to generate sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length from said one or more nucleic acid molecule(s) and said extraction unit is configured to compared said nucleotide fragments of “n+1”, “n+2” . . . “n+x” length to the nucleic acid sequences in a database.
38 . The system of claim 36 , wherein said nucleotide fragment of “n” length is compared to said reference nucleic acid sequences in real time, or as fast as said fragment of “n” length is generated, while the sequencing unit continues to generate sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length from said one or more nucleic acid molecule(s), and the processing unit compares said sequence information of “n+1”, “n+2” . . . “n+x” nucleotide fragments in length, in real-time, or as fast as the fragments are generated to the nucleic acid sequences in a database.
39 . A method of identifying a biological material in a sample, comprising:
(i) obtaining a sample comprising said biological material, (ii) extracting one or more nucleic acid molecule(s) from said sample, (iii) generating sequence information, comprising a sequence of a nucleotide fragment from said one or more nucleic acid molecule(s), (iv) comparing said sequence of a nucleotide fragment to nucleic acid sequences in a database;
and if said comparison of said sequence of a nucleotide fragment does not result in a match identifying the biological material in said sample, then the method further comprises:
(v) generating additional sequence information from said one or more nucleic acid molecule(s), wherein said additional sequence information comprises a sequence of a nucleotide fragment consisting of one additional nucleotide,
(vi) comparing said additional sequence information to nucleic acid sequences in a database immediately following the generation of said additional sequence information,
and repeating steps (v)-(vi) until a match results in the identification of the biological material is said sample.
40 . A method of identifying a biological material in a sample, comprising:
(i) obtaining a sample comprising said biological material, (ii) extracting one or more nucleic acid molecule(s) from said sample, (iii) amplifying said one or more nucleic acid molecule(s) to yield a plurality of one or more nucleic acid molecule(s), (iii) generating a plurality of sequence information, comprising a plurality of sequences of a nucleotide fragment, from said plurality of one or more nucleic acid molecule(s), (iv) comparing said plurality of sequences of a nucleotide fragment to nucleic acid sequences in a database,
and if said comparison of said plurality of sequences of a nucleotide fragment does not result in a match identifying the biological material in said sample, then the method further comprises:
(v) generating plurality of additional sequence information from said one or more nucleic acid molecule(s), wherein said additional sequence information comprises a sequence of a nucleotide fragment consisting of one additional nucleotide,
(vi) comparing said additional sequence information to nucleic acid sequences in a database immediately following the generation of said additional sequence information,
and repeating steps (v)-(vi) until a match results in the identification of the biological material is said sample.
41 . The methods of claim 39 or 40 , wherein the comparison to the nucleic acid sequences in a database is performed via probabilistic matching as fast as the sequence information is generated.Join the waitlist — get patent alerts
Track US2019295687A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.