US2011295864A1PendingUtilityA1

Iterative fact-extraction

Assignee: BETZ MARTINPriority: May 29, 2010Filed: May 29, 2010Published: Dec 1, 2011
Est. expiryMay 29, 2030(~3.8 yrs left)· nominal 20-yr term from priority
G06F 16/38G06F 16/383
34
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Some embodiments provide a method for identifying a first pattern formed by a first set of document elements. The method associates a tag that identifies the first pattern with the document. The method then identifies a second pattern formed by a second set of document elements and the tag. The method associates a second tag with the document. The second tag identifies the second pattern and is associated with a fact. Some embodiments provide a graphical user interface (GUI) for manually processing tags associated with a document. Further, different embodiments provide a system for performing iterative fact-extraction on a set of documents.

Claims

exact text as granted — not AI-modified
1 . A machine-implemented method for extracting a fact from a document, the document comprising a plurality of document elements, the method comprising:
 identifying a first pattern formed by a first set of document elements;   associating a first tag with the document, the first tag identifying the first pattern;   identifying a second pattern formed by a second set of document elements and the first tag;   associating a second tag with the document; and   based on the second tag, recording a first fact for the document.   
     
     
         2 . The method of  claim 1 , wherein the first pattern is defined by a list including known words and expressions. 
     
     
         3 . The method of  claim 1 , wherein the second pattern is defined by a rule that specifies a required relationship between the second set of document elements and the first tag. 
     
     
         4 . The method of  claim 1 , wherein the first tag identifies the first pattern as a company, the second tag identifies the second pattern as an action verb, and the first fact is a new company hire. 
     
     
         5 . The method of  claim 1 , wherein the first tag identifies the first pattern as a person, the second tag identifies the second pattern as a quote, and the first fact is the quote attributed to the person. 
     
     
         6 . The method of  claim 1 , wherein the first tag identifies the first pattern as a person, the second tag identifies the second pattern as a gender pronoun, and the first fact is the person being male or female. 
     
     
         7 . The method of  claim 1  further comprising:
 identifying a third pattern formed by a third set of document elements, the first tag, and the second tag; and 
 associating a third tag with the document, the third tag identifying the third pattern and associated with a second fact. 
 
     
     
         8 . A computer readable storage medium including a computer program, the computer program including instructions for providing a graphical user interface (GUI) for manually processing tags associated with a document, the GUI comprising:
 a first UI item for selecting a script for performing iterative fact-extraction on text data from the document;   a text box UI item for inputting the text data from the document for the iterative fact-extraction;   a first display portion for presenting identified patterns from the inputted text data resulting from the iterative fact-extraction; and   a second display portion for providing a plurality of UI items representing a plurality of tags associated with the identified patterns.   
     
     
         9 . The computer readable storage medium of  claim 8 , wherein the plurality of UI items allows a user to modify a first tag associated with a particular pattern to a second tag when the user selects a second UI item that represents the second tag. 
     
     
         10 . The computer readable storage medium of  claim 8 , wherein the plurality of UI items further includes a second UI item for removing a tag associated with a particular pattern. 
     
     
         11 . A system for performing iterative fact-extraction on a set of documents, the system comprising:
 a pattern analysis engine for identifying a set of patterns in the set of documents; and   a tag engine for annotating the set of documents with respective tags that are associated with facts in the set of documents.   
     
     
         12 . The system of  claim 11 , wherein the pattern analysis engine executes a set of pattern analysis instructions to identify the set of patterns, said set of pattern analysis instructions defining the set of patterns to identify. 
     
     
         13 . The system of  claim 11 , wherein the set of documents are stored in a document storage. 
     
     
         14 . The system of  claim 11 , wherein the tags are stored in a tag storage. 
     
     
         15 . The system of  claim 14  further comprising:
 a fact processing module for processing the stored tags to extract a set of facts associated with the tags; and 
 a query processor for executing search queries on the set of facts to retrieve facts that match the search queries. 
 
     
     
         16 . The system of  claim 11  further comprising a document crawler module for communicating with a network to retrieve the set of documents on a real-time basis. 
     
     
         17 . The system of  claim 11  further comprising a document crawler module for communicating with a network to retrieve the set of documents on a periodic basis. 
     
     
         18 . The system of  claim 12  further comprising a file handler module for receiving scripts that are embedded with the set of pattern analysis instructions. 
     
     
         19 . The computer readable medium of  claim 8 , wherein the first display portion provides a second UI item for editing the identified patterns. 
     
     
         20 . The computer readable medium of  claim 8 , wherein the first display portion provides a second UI item for removing an identified pattern and any tag associated with the pattern.

Join the waitlist — get patent alerts

Track US2011295864A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.