Ingestion planning for complex tables
Abstract
Embodiments of the present invention disclose a method, computer program product, and system for generating a plan for document processing. A plurality of electronic documents are received, by a computer, using a network. The plurality of electronic documents are analyzed, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents. Textual data is identified within the identified tabular data, of the analyzed plurality of electronic documents. Textual hints are generated, based on the identified textual data within the identified tabular data. References are identified, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents. A count of references is calculated, associated with one or more sets of tabular data. A priority score is calculated based on the count of references, and an ingestion plan is generated, based on the calculated priority score.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for generating a plan for document processing, the method comprising:
receiving a plurality of electronic documents, by a computer using a network; analyzing the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents; identifying textual data within the identified tabular data, of the analyzed plurality of electronic documents; generating textual hints, based on the identified textual data within the identified tabular data; identifying references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents; calculating a count of references associated with one or more sets of tabular data; calculating a priority score based on the count of references; and generating an ingestion plan, based on the calculated priority score.
2 . The method of claim 1 , further comprising:
communicating the generated ingestion plan; and executing the generated ingestion plan.
3 . The method of claim 1 , wherein calculating a count of references further comprises:
generating an index of textual hints; and generating a lookup set based on the generated index of textual hints.
4 . The method of claim 1 , wherein generating an ingestion plan further comprises:
generating a ordered list of the one or more subsets of the identified tabular data, based on the calculated priority score associated with each of the one or more subsets of identified tabular data; communicating a subset of the generated list, based on a threshold score, for display; and communicating the subset of the generated list for further processing.
5 . The method of claim 1 , further comprising:
communicating an importance value option for display to a user; receiving an importance value, based on receiving a user selection of the importance value option; modifying the calculated priority score, based on the received importance value; and generating a second list of one or more subsets of identified tabular data, based on the modified calculated priority score.
6 . The method of claim 1 , wherein identifying textual data is performed by a natural language analysis.
7 . The method of claim 1 , further comprising:
calculating a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and modifying the calculated priority score, based on the calculated percent value.
8 . The method of claim 1 , further comprising:
processing the generated list of one or more subsets of tabular data, wherein the processing includes:
ordering the one or more subsets of the generated list, based on the associated calculated priority score;
extracting data associated with the one or more subsets of tabular data;
annotating the one or more subsets of tabular data; and
extracting data from the annotated one or more subsets of tabular data.
9 . The method of claim 8 , wherein the extracting data associated with the one or more subsets of tabular data occurs subsequently to one or more of: an Extensible Markup Language formatting; an Unstructured Information Management Architecture formatting; and an OpenDocument formatting.
10 . A computer program product for generating a plan for document processing, the computer program product comprising:
one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising:
instructions to instructions to receive a plurality of electronic documents, by a computer using a network;
instructions to analyze the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents;
instructions to identify textual data within the identified tabular data, of the analyzed plurality of electronic documents;
instructions to generate textual hints, based on the identified textual data within the identified tabular data;
instructions to identify references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents;
instructions to calculate a count of references associated with one or more sets of tabular data;
instructions to calculate a priority score based on the count of references; and
instructions to generate an ingestion plan, based on the calculated priority score.
11 . The computer program product of claim 10 , further comprising:
instructions to communicate the generated ingestion plan; and instructions to execute the generated ingestion plan.
12 . The computer program product of claim 10 , wherein calculating a count of references further comprises:
instructions to generate an index of textual hints; and instructions to generate a lookup set based on the generated index of textual hints.
13 . The computer program product of claim 10 , wherein generating an ingestion plan further comprises:
instructions to generate a ordered list of the one or more subsets of the identified tabular data, based on the calculated priority score associated with each of the one or more subsets of identified tabular data; instructions to communicate a subset of the generated list, based on a threshold score, for display; and instructions to communicate the subset of the generated list for further processing.
14 . The computer program product of claim 10 , further comprising:
instructions to communicate an importance value option for display to a user; instructions to receive an importance value, based on receiving a user selection of the importance value option; instructions to modify the calculated priority score, based on the received importance value; and instructions to generate a second list of one or more subsets of identified tabular data, based on the modified calculated priority score.
15 . The computer program product of claim 10 , wherein instructions to identify textual data is performed by a natural language analysis.
16 . The computer program product of claim 10 , further comprising:
instructions to calculate a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and instructions to modify the calculated priority score, based on the calculated percent value.
17 . The computer program product of claim 10 , further comprising:
instructions to process the generated list of one or more subsets of tabular data, wherein the instructions to process include:
instructions to order the one or more subsets of the generated list, based on the associated calculated priority score;
instructions to extract data associated with the one or more subsets of tabular data;
instructions to annotate the one or more subsets of tabular data; and
instructions to extract data from the annotated one or more subsets of tabular data.
18 . The computer program product of claim 17 , wherein the instructions to extract data associated with the one or more subsets of tabular data occurs subsequently to one or more of: instructions to Extensible Markup Language format; instructions to Unstructured Information Management Architecture format; and instructions to OpenDocument format.
19 . A computer system for generating a plan for document processing, the computer system comprising:
one or more computer processors; one or more computer-readable storage media; program instructions stored on the computer-readable storage media for execution by at least one of the one or more processors, the program instructions comprising:
instructions to receive a plurality of electronic documents, by a computer using a network;
instructions to analyze the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents;
instructions to identify textual data within the identified tabular data, of the analyzed plurality of electronic documents;
instructions to generate textual hints, based on the identified textual data within the identified tabular data;
instructions to identify references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents;
instructions to calculate a count of references associated with one or more sets of tabular data;
instructions to calculate a priority score based on the count of references; and
instructions to generate an ingestion plan, based on the calculated priority score.
20 . The system of claim 19 , further comprising:
instructions to calculate a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and instructions to modify the calculated priority score, based on the calculated percent value.Join the waitlist — get patent alerts
Track US2017116194A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.