US2017098034A1PendingUtilityA1

Constructing custom knowledgebases and sequence datasets with publications

Assignee: BATTELLE MEMORIAL INSTITUTEPriority: May 16, 2014Filed: Dec 21, 2016Published: Apr 6, 2017
Est. expiryMay 16, 2034(~7.8 yrs left)· nominal 20-yr term from priority
G06F 19/24G06N 5/022G06F 19/28G16B 50/30G16B 30/00G16B 40/00G16B 50/00
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Illustrative embodiments of custom knowledgebases and sequence datasets, as well as related methods, are disclosed. In one illustrative embodiment, one or more computer-readable media may comprise a custom knowledgebase and an associated sequence dataset. The custom knowledgebase may comprise a plurality of assertions that have been automatically extracted from a plurality of publications, where each of the plurality of assertions encodes a relationship between a subject and an object. The sequence dataset may comprise a plurality of called biological sequences, where each of the plurality of called biological sequences is associated with one or more of the plurality of assertions of the custom knowledgebase.

Claims

exact text as granted — not AI-modified
1 - 20 . (canceled) 
     
     
         21 . A method comprising:
 automatically extracting a plurality of assertions from a plurality of publications, wherein each of the plurality of assertions encodes a relationship between a subject and an object;   automatically extracting one or more called biological sequences from the plurality of publications;   extracting additional called biological sequences from one or more publicly available databases;   grouping the additional called biological sequences with the one or more called biological sequences automatically extracted from the plurality of publications in response to one or more resemblance criteria being met; and   associating each group of called biological sequences with one or more of the plurality of assertions.   
     
     
         22 . The method of  claim 21 , wherein each of the one or more resemblance criteria is predetermined. 
     
     
         23 . The method of  claim 21 , wherein automatically extracting the plurality of assertions from the plurality of publications comprises utilizing natural language processing software to derive the plurality of assertions from the text of the plurality of publications. 
     
     
         24 . The method of  claim 23 , wherein the plurality of publications comprises peer-reviewed articles selected by subject matter experts in a field associated with the peer-reviewed articles. 
     
     
         25 . The method of  claim 23 , wherein the natural language processing software has been trained by subject matter experts in a field associated with the plurality of publications to recognize relevant assertions in the text of the plurality of publications. 
     
     
         26 . The method of  claim 23 , wherein each of the plurality of assertions is expressed as a Resource Description Framework (RDF) triple. 
     
     
         27 . The method of  claim 21 , further comprising manually editing the associations between each group of called biological sequences and the plurality by subject matter experts in a field associated with the plurality of publications. 
     
     
         28 . One or more tangible non-transitory computer-readable media comprising a plurality of instructions that, when executed by computing device, causes the computing device to:
 automatically extract a plurality of assertions from a plurality of publications, wherein each of the plurality of assertions encodes a relationship between a subject and an object;   automatically extract one or more called biological sequences from the plurality of publications;   extract additional called biological sequences from one or more publicly available databases;   group the additional called biological sequences with the one or more called biological sequences automatically extracted from the plurality of publications in response to one or more resemblance criteria being met; and   associate each group of called biological sequences with one or more of the plurality of assertions.   
     
     
         29 . The one or more tangible non-transitory computer-readable media of  claim 28 , wherein each of the one or more resemblance criteria is predetermined. 
     
     
         30 . The one or more tangible non-transitory computer-readable media of  claim 28 , wherein to automatically extract the plurality of assertions from the plurality of publications comprises to utilize natural language processing software to derive the plurality of assertions from the text of the plurality of publications. 
     
     
         31 . The one or more tangible non-transitory computer-readable media of  claim 30 , wherein the plurality of publications comprises peer-reviewed articles selected by subject matter experts in a field associated with the peer-reviewed articles. 
     
     
         32 . The one or more tangible non-transitory computer-readable media of  claim 30 , wherein the natural language processing software has been trained by subject matter experts in a field associated with the plurality of publications to recognize relevant assertions in the text of the plurality of publications. 
     
     
         33 . The one or more computer-readable media of  claim 30 , wherein each of the plurality of assertions is expressed as a Resource Description Framework (RDF) triple. 
     
     
         34 . A method comprising:
 comparing a plurality of sample biological sequences to a plurality of called biological sequences included in a sequence dataset;   retrieving, from a custom knowledgebase associated with the sequence dataset, one or more assertions that are associated with a called biological sequence of the sequence dataset that resembles one of the plurality of sample biological sequences, wherein the one of the plurality of sample biological sequences is not in the sequence dataset; and   determining one or more probable characteristics associated with the sample biological sequence that resembles the called biological sequence of the sequence dataset using the one or more assertions retrieved from the custom knowledgebase.   
     
     
         35 . The method of  claim 34 , further comprising generating the plurality of sample biological sequences using massively parallel sequencing of a metagenomic sample. 
     
     
         36 . The method of  claim 34 , wherein determining one or more probable characteristics associated with the sample biological sequence comprises determining one or more antibiotics likely to be resisted. 
     
     
         37 . The method of  claim 36 , further comprising generating a report that comprises a ranked listing of the antibiotics likely to be resisted. 
     
     
         38 . The method of  claim 34 , wherein each of the one or more assertions is expressed as a Resource Description Framework (RDF) triple. 
     
     
         39 . The method of  claim 34 , wherein the of called biological sequences of the sequence dataset comprise at least one of called biological sequences that provide resistance to one or more antibiotics and called biological sequences that mediate regulation of antibiotic resistance. 
     
     
         40 . The method of  claim 39 , wherein the one or more assertions of the custom knowledgebase comprise assertions that encode relationships between the called biological sequences of the sequence dataset and at least one of antibiotic resistance elements and regulatory elements.

Join the waitlist — get patent alerts

Track US2017098034A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.