US2008126080A1PendingUtilityA1

Creation of structured data from plain text

Assignee: ARIBA INCPriority: Jan 8, 2001Filed: Oct 31, 2007Published: May 29, 2008
Est. expiryJan 8, 2021(expired)· nominal 20-yr term from priority
Y10S707/99936G06F 40/211Y10S707/99942G06F 40/131G06F 16/258G06F 40/30Y10S707/99943G06F 40/143
36
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and system for converting plain text into structured data. Parse trees for the plain text are generated based on the grammar of a natural language, the parse trees are mapped on to instance trees generated based on an application-specific model. The best map is chosen, and the instance tree is passing to an application for execution. The method and system can be used both for populating a database and/or for retrieving data from a database based on a query.

Claims

exact text as granted — not AI-modified
1 - 22 . (canceled) 
   
   
       23 . A system, comprising:
 a processor configured to:
 tokenize a plain text description; 
 create parse trees from the tokenized plain text description based on grammar from a grammar storage area; 
 generate an instance tree from each parse tree based upon an application domain specific natural markup language provided by a natural markup language model module; 
 discard each invalid or incomplete instance tree; 
 choose an instance tree from remaining instance trees representing a best map based upon a cost function; and 
 process the best map with a domain markup language generator to generate a structured data representation; and 
   a memory coupled to the processor and configured to provide the processor with instructions.   
   
   
       24 . The system of  claim 23  wherein the processor is further configured to use the structured data representation to populate a database. 
   
   
       25 . The system of  claim 23  wherein the processor is further configured to use the structured data representation to query a database. 
   
   
       26 . The system of  claim 23  wherein the processor is further configured to use the structured data representation to invoke an application. 
   
   
       27 . The system of  claim 23 , wherein the cost function comprises choosing maps with less structure over maps with more created structure. 
   
   
       28 . The system of  claim 23 , wherein the cost function comprises:
 choosing maps that use the most tokens contained in compact groups over maps using fewer tokens spread further over text segments;   choosing maps with the tightest possible bindings; and   choosing maps that have fewer objects.   
   
   
       29 . The system of  claim 23 , wherein the cost function comprises:
 (a) choosing a map with the most tokens;   (b) if maps are equal under (a), then choosing a map having a topmost expression farthest from a root of the map;   (c) if maps are equal under (a) and (b), then choosing a map with a least distance between tokens;   (d) if maps are equal under (a) through (c), then choosing a map with fewer objects created by enumerations;   (e) if maps are equal under (a) through (d), then choosing a map with fewer unused primitives;   (f) if maps are equal under (a) through (e), then choosing a map with fewer objects created by database lookup;   (g) if maps are equal under (a) through (f), then choosing a map with fewer natural markup language objects;   (h) if maps are equal under (a) through (g), then choosing a map with fewer inferred objects.   
   
   
       30 . The system of  claim 29 , wherein the cost function further comprises:
 (i) if maps are equal under (a) through (h), then regarding all maps as equally valid.   
   
   
       31 . The system of  claim 23 , wherein all possible parse trees from the tokenized plain text are created. 
   
   
       32 . The system of  claim 23 , wherein the processor is further configured to represent all of the parse trees in a single directed acyclic graph. 
   
   
       33 . The system of  claim 23 , wherein the grammar from the grammar storage area is context free. 
   
   
       34 . A computer program product embodied in a computer readable medium and comprising computer instructions for:
 tokenizing a plain text description;   creating parse trees from the tokenized plain text description based on grammar from a grammar storage area;   generating an instance tree from each parse tree based upon an application domain specific natural markup language provided by a natural markup language model module;   discarding each invalid or incomplete instance tree;   choosing an instance tree from remaining instance trees representing a best map based upon a cost function; and   processing the best map with a domain markup language generator to generate a structured data representation.   
   
   
       35 . The computer program product of  claim 34  further comprising computer instructions for using the structured data representation to populate a database. 
   
   
       36 . The computer program product of  claim 34  further comprising computer instructions for using the structured data representation to query a database. 
   
   
       37 . The computer program product of  claim 34  further comprising computer instructions for using the structured data representation to invoke an application. 
   
   
       38 . The computer program product of  claim 34 , wherein the cost function comprises choosing maps with less structure over maps with more created structure. 
   
   
       39 . The computer program product of  claim 34 , wherein the cost function comprises:
 choosing maps that use the most tokens contained in compact groups over maps using fewer tokens spread further over text segments;   choosing maps with the tightest possible bindings; and   choosing maps that have fewer objects.   
   
   
       40 . The computer program product of  claim 34 , wherein the cost function comprises:
 (a) choosing a map with the most tokens;   (b) if maps are equal under (a), then choosing a map having a topmost expression farthest from a root of the map;   (c) if maps are equal under (a) and (b), then choosing a map with a least distance between tokens;   (d) if maps are equal under (a) through (c), then choosing a map with fewer objects created by enumerations;   (e) if maps are equal under (a) through (d), then choosing a map with fewer unused primitives;   (f) if maps are equal under (a) through (e), then choosing a map with fewer objects created by database lookup;   (g) if maps are equal under (a) through (f), then choosing a map with fewer natural markup language objects;   (h) if maps are equal under (a) through (g), then choosing a map with fewer inferred objects.   
   
   
       41 . The computer program product of  claim 40 , wherein the cost function further comprises:
 (i) if maps are equal under (a) through (h), then regarding all maps as equally valid.   
   
   
       42 . The computer program product of  claim 34 , wherein the grammar from the grammar storage area is context free.

Join the waitlist — get patent alerts

Track US2008126080A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.