System and method for extracting data from a source
Abstract
Systems and methods are provided for extracting information from a document. An exemplary method includes: receiving a plurality of sample documents; retrieving image information for the sample documents; converting each of the image information associated with the plurality of sample documents into a knowledge graph; determining a plurality of rules applicable for processing the sample documents; retrieving the rules applicable for processing the sample documents; parsing the rules to create a plurality of abstract syntax trees; generating a plurality of rule results for the rules applicable for processing the sample documents; applying the rule results to a probabilistic model to determine accuracy of the rule results; generating a plurality of scores for each of the rules indicating accuracy of the rules; determining whether the scores require updating the probabilistic model; and updating the probabilistic model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for extracting information from documents, comprising:
receiving, using one or more processors, a plurality of sample documents; retrieving, using the one or more processors, image information for the sample documents; generating, using the one or more processors, hierarchical and spatial data of the image information of the sample documents to be used in a knowledge graph, wherein the knowledge graph defines hierarchical and spatial relationships between the sample documents; determining, using the one or more processors, a plurality of rules applicable for processing the sample documents; upon determining the rules applicable for processing the sample documents, retrieving, using the one or more processors, the rules applicable for processing the sample documents; parsing, using the one or more processors, the rules to create a plurality of abstract syntax trees; generating, using the one or more processors, the abstract syntax trees, and the knowledge graph, a plurality of rule results for the rules applicable for processing the sample documents, wherein each of the abstract syntax trees is evaluated on the knowledge graph to produce the rule results; applying, using the one or more processors, the rule results to a probabilistic model to determine accuracy of the rule results; generating, using the one or more processors and the probabilistic model, a plurality of scores for each of the rules indicating accuracy of the rules; determining, using the one or more processors, whether the scores require updating the probabilistic model, and upon determining the probabilistic model requires updating, updating, using the one or more processors, the probabilistic model.
2 . The method of claim 1 , wherein retrieving the image information comprises performing an optical character recognition (OCR) operation on the sample documents.
3 . The method of claim 1 , retrieving the image information comprises adapting to spatial noise introduced by an optical character recognition (OCR) operation on the sample documents.
4 . The method of claim 1 , wherein converting each of the image information comprises determining spatial and textual information of the sample documents to produce the knowledge graph.
5 . The method of claim 1 , wherein converting each of the image information comprises adapting to spatial noise introduced by the knowledge graph.
6 . The method of claim 1 , wherein determining the plurality of rules applicable for processing comprises determine whether the plurality of rules are existing rules or rules under development.
7 . The method of claim 6 , wherein the plurality of rules are formed using a domain specific language.
8 . The method of claim 1 , wherein parsing each of the rules comprises utilizing a recurrent neural network to parse each of the rules.
9 . The method of claim 1 , wherein applying the rule results to the probabilistic model comprises utilizing a Naïve Bayesian methodology to assess the rule results.
10 . The method of claim 1 , wherein applying the rule results to the probabilistic model comprises allowing the plurality of rules to dynamically adapt to each instance of at least one of the sample documents being processed.
11 . The method of claim 1 , wherein determining whether the scores require updating comprises identifying the rules having top scores.
12 . The method of claim 1 , wherein determining whether the scores require updating comprises determining whether the probabilistic model comprises a specific shape.
13 . The method of claim 1 , wherein determining whether the scores are sufficient comprises iteratively growing the plurality of rules to improve the probabilistic model.
14 . A system for extracting information from a document, the system comprising
one or more computing device processors; and one or more computing device memories, coupled to the one or more computing device processors, the one or more computing device memories storing instructions executed by the one or more computing device processors, wherein the instructions are configured to:
receive a plurality of sample documents;
retrieve image information for the sample documents;
generate hierarchical and spatial data of the image information of the sample documents to be used in a knowledge graph, wherein the knowledge graph defines hierarchical and spatial relationships between the sample documents;
determine a plurality of rules applicable for processing the sample documents;
upon determining the rules applicable for processing the sample documents, retrieve the rules applicable for processing the sample documents;
parse the rules to create a plurality of abstract syntax trees;
generate, using the abstract syntax trees and the knowledge graph, a plurality of rule results for the rules applicable for processing the sample documents, wherein each of the abstract syntax trees is evaluated on the knowledge graph to produce the rule results;
apply the rule results to a probabilistic model to determine accuracy of the rule results;
generate, using the probabilistic model, a plurality of scores for each of the rules indicating accuracy of the rules;
determine whether the scores require updating the probabilistic model; and
upon determining the probabilistic model requires updating, update the probabilistic model.
15 . The system of claim 14 , wherein the instructions are further configured to provide a user interface for displaying the rules that have executed.
16 . The system of claim 14 , wherein the instructions are further configured to provide a user interface for displaying details of at least one of the sample documents.
17 . The system of claim 14 , wherein the instructions are further configured to provide a user interface for allowing a user to input at least one of the sample documents.
18 . The system of claim 14 , wherein the instructions are further configured to provide a user interface for displaying details of at least one of the rules that is actively being developed.
19 . The system of claim 14 , wherein the instructions are further configured to provide a user interface for displaying details of the scores for the plurality of rules.
20 . A method for extraction information from documents comprising:
receiving, using one or more processors, a plurality of sample documents; retrieving, using the one or more processors, image information for the sample documents; generating, using the one or more processors, hierarchical and spatial data of the image information of the sample documents to be used in a knowledge graph, wherein the knowledge graph defines hierarchical and spatial relationships between the sample documents; determining, using the one or more processors, a plurality of rules applicable for processing the sample documents; upon determining the rules applicable for processing the sample documents, retrieving, using the one or more processors, the rules applicable for processing the sample documents; parsing, using the one or more processors, the rules to create a plurality of abstract syntax trees; generating, using the one or more processors, the abstract syntax trees, and the knowledge graph, a plurality of rule results for the rules applicable for processing the sample documents, wherein each of the abstract syntax trees is evaluated on the knowledge graph to produce the rule results; applying, using the one or more processors, the rule results to a probabilistic model to determine accuracy of the rule results; generating, using the one or more processors and the probabilistic model, a plurality of scores for each of the rules indicating accuracy of the rules; determining, using the one or more processors, whether there is at least one additional sample document or at least one additional rule added by a user; upon determining the user added the least one additional sample document, updating, using the one or more processors, the knowledge graph to reflect changes presented by the at least one additional sample document; upon determining the user added the least one additional rule, updating, using the one or more processors, the abstract index trees to reflect changes presented by the at least one additional rule; updating, using the one or more processors, the updated knowledge graph, and the updated abstract syntax trees, the scores for each of the rules and the at least one additional rule; determining, using the one or more processors, whether the updated scores require updating the probabilistic model, and upon determining the probabilistic model requires updating, updating, using the one or more processors, the probabilistic model.Join the waitlist — get patent alerts
Track US2023260307A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.