Iterative fact-extraction
Abstract
Some embodiments provide a method for identifying a first pattern formed by a first set of document elements. The method associates a tag that identifies the first pattern with the document. The method then identifies a second pattern formed by a second set of document elements and the tag. The method associates a second tag with the document. The second tag identifies the second pattern and is associated with a fact. Some embodiments provide a graphical user interface (GUI) for manually processing tags associated with a document. Further, different embodiments provide a system for performing iterative fact-extraction on a set of documents.
Claims
exact text as granted — not AI-modified1 . A machine-implemented method for extracting a fact from a document, the document comprising a plurality of document elements, the method comprising:
identifying a first pattern formed by a first set of document elements; associating a first tag with the document, the first tag identifying the first pattern; identifying a second pattern formed by a second set of document elements and the first tag; associating a second tag with the document; and based on the second tag, recording a first fact for the document.
2 . The method of claim 1 , wherein the first pattern is defined by a list including known words and expressions.
3 . The method of claim 1 , wherein the second pattern is defined by a rule that specifies a required relationship between the second set of document elements and the first tag.
4 . The method of claim 1 , wherein the first tag identifies the first pattern as a company, the second tag identifies the second pattern as an action verb, and the first fact is a new company hire.
5 . The method of claim 1 , wherein the first tag identifies the first pattern as a person, the second tag identifies the second pattern as a quote, and the first fact is the quote attributed to the person.
6 . The method of claim 1 , wherein the first tag identifies the first pattern as a person, the second tag identifies the second pattern as a gender pronoun, and the first fact is the person being male or female.
7 . The method of claim 1 further comprising:
identifying a third pattern formed by a third set of document elements, the first tag, and the second tag; and
associating a third tag with the document, the third tag identifying the third pattern and associated with a second fact.
8 . A computer readable storage medium including a computer program, the computer program including instructions for providing a graphical user interface (GUI) for manually processing tags associated with a document, the GUI comprising:
a first UI item for selecting a script for performing iterative fact-extraction on text data from the document; a text box UI item for inputting the text data from the document for the iterative fact-extraction; a first display portion for presenting identified patterns from the inputted text data resulting from the iterative fact-extraction; and a second display portion for providing a plurality of UI items representing a plurality of tags associated with the identified patterns.
9 . The computer readable storage medium of claim 8 , wherein the plurality of UI items allows a user to modify a first tag associated with a particular pattern to a second tag when the user selects a second UI item that represents the second tag.
10 . The computer readable storage medium of claim 8 , wherein the plurality of UI items further includes a second UI item for removing a tag associated with a particular pattern.
11 . A system for performing iterative fact-extraction on a set of documents, the system comprising:
a pattern analysis engine for identifying a set of patterns in the set of documents; and a tag engine for annotating the set of documents with respective tags that are associated with facts in the set of documents.
12 . The system of claim 11 , wherein the pattern analysis engine executes a set of pattern analysis instructions to identify the set of patterns, said set of pattern analysis instructions defining the set of patterns to identify.
13 . The system of claim 11 , wherein the set of documents are stored in a document storage.
14 . The system of claim 11 , wherein the tags are stored in a tag storage.
15 . The system of claim 14 further comprising:
a fact processing module for processing the stored tags to extract a set of facts associated with the tags; and
a query processor for executing search queries on the set of facts to retrieve facts that match the search queries.
16 . The system of claim 11 further comprising a document crawler module for communicating with a network to retrieve the set of documents on a real-time basis.
17 . The system of claim 11 further comprising a document crawler module for communicating with a network to retrieve the set of documents on a periodic basis.
18 . The system of claim 12 further comprising a file handler module for receiving scripts that are embedded with the set of pattern analysis instructions.
19 . The computer readable medium of claim 8 , wherein the first display portion provides a second UI item for editing the identified patterns.
20 . The computer readable medium of claim 8 , wherein the first display portion provides a second UI item for removing an identified pattern and any tag associated with the pattern.Join the waitlist — get patent alerts
Track US2011295864A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.