US2017116194A1PendingUtilityA1

Ingestion planning for complex tables

Assignee: IBMPriority: Oct 23, 2015Filed: Oct 23, 2015Published: Apr 27, 2017
Est. expiryOct 23, 2035(~9.2 yrs left)· nominal 20-yr term from priority
G06F 16/93G06F 16/3329G06F 16/2282G06F 16/48G06F 16/148G06F 16/313G06F 16/254G06F 40/56G06F 16/41G06F 16/243G06F 40/151G06F 16/3326G06F 16/24578G06F 16/3344G06F 40/177G06F 40/20G06F 16/345G06F 17/30011G06F 17/30684G06F 17/3053G06F 17/30339
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Embodiments of the present invention disclose a method, computer program product, and system for generating a plan for document processing. A plurality of electronic documents are received, by a computer, using a network. The plurality of electronic documents are analyzed, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents. Textual data is identified within the identified tabular data, of the analyzed plurality of electronic documents. Textual hints are generated, based on the identified textual data within the identified tabular data. References are identified, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents. A count of references is calculated, associated with one or more sets of tabular data. A priority score is calculated based on the count of references, and an ingestion plan is generated, based on the calculated priority score.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for generating a plan for document processing, the method comprising:
 receiving a plurality of electronic documents, by a computer using a network;   analyzing the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents;   identifying textual data within the identified tabular data, of the analyzed plurality of electronic documents;   generating textual hints, based on the identified textual data within the identified tabular data;   identifying references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents;   calculating a count of references associated with one or more sets of tabular data;   calculating a priority score based on the count of references; and   generating an ingestion plan, based on the calculated priority score.   
     
     
         2 . The method of  claim 1 , further comprising:
 communicating the generated ingestion plan; and   executing the generated ingestion plan.   
     
     
         3 . The method of  claim 1 , wherein calculating a count of references further comprises:
 generating an index of textual hints; and   generating a lookup set based on the generated index of textual hints.   
     
     
         4 . The method of  claim 1 , wherein generating an ingestion plan further comprises:
 generating a ordered list of the one or more subsets of the identified tabular data, based on the calculated priority score associated with each of the one or more subsets of identified tabular data;   communicating a subset of the generated list, based on a threshold score, for display; and   communicating the subset of the generated list for further processing.   
     
     
         5 . The method of  claim 1 , further comprising:
 communicating an importance value option for display to a user;   receiving an importance value, based on receiving a user selection of the importance value option;   modifying the calculated priority score, based on the received importance value; and   generating a second list of one or more subsets of identified tabular data, based on the modified calculated priority score.   
     
     
         6 . The method of  claim 1 , wherein identifying textual data is performed by a natural language analysis. 
     
     
         7 . The method of  claim 1 , further comprising:
 calculating a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and   modifying the calculated priority score, based on the calculated percent value.   
     
     
         8 . The method of  claim 1 , further comprising:
 processing the generated list of one or more subsets of tabular data, wherein the processing includes:
 ordering the one or more subsets of the generated list, based on the associated calculated priority score; 
 extracting data associated with the one or more subsets of tabular data; 
 annotating the one or more subsets of tabular data; and 
 extracting data from the annotated one or more subsets of tabular data. 
   
     
     
         9 . The method of  claim 8 , wherein the extracting data associated with the one or more subsets of tabular data occurs subsequently to one or more of: an Extensible Markup Language formatting; an Unstructured Information Management Architecture formatting; and an OpenDocument formatting. 
     
     
         10 . A computer program product for generating a plan for document processing, the computer program product comprising:
 one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising:
 instructions to instructions to receive a plurality of electronic documents, by a computer using a network; 
 instructions to analyze the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents; 
 instructions to identify textual data within the identified tabular data, of the analyzed plurality of electronic documents; 
 instructions to generate textual hints, based on the identified textual data within the identified tabular data; 
 instructions to identify references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents; 
 instructions to calculate a count of references associated with one or more sets of tabular data; 
 instructions to calculate a priority score based on the count of references; and 
 instructions to generate an ingestion plan, based on the calculated priority score. 
   
     
     
         11 . The computer program product of  claim 10 , further comprising:
 instructions to communicate the generated ingestion plan; and   instructions to execute the generated ingestion plan.   
     
     
         12 . The computer program product of  claim 10 , wherein calculating a count of references further comprises:
 instructions to generate an index of textual hints; and   instructions to generate a lookup set based on the generated index of textual hints.   
     
     
         13 . The computer program product of  claim 10 , wherein generating an ingestion plan further comprises:
 instructions to generate a ordered list of the one or more subsets of the identified tabular data, based on the calculated priority score associated with each of the one or more subsets of identified tabular data;   instructions to communicate a subset of the generated list, based on a threshold score, for display; and   instructions to communicate the subset of the generated list for further processing.   
     
     
         14 . The computer program product of  claim 10 , further comprising:
 instructions to communicate an importance value option for display to a user;   instructions to receive an importance value, based on receiving a user selection of the importance value option;   instructions to modify the calculated priority score, based on the received importance value; and   instructions to generate a second list of one or more subsets of identified tabular data, based on the modified calculated priority score.   
     
     
         15 . The computer program product of  claim 10 , wherein instructions to identify textual data is performed by a natural language analysis. 
     
     
         16 . The computer program product of  claim 10 , further comprising:
 instructions to calculate a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and   instructions to modify the calculated priority score, based on the calculated percent value.   
     
     
         17 . The computer program product of  claim 10 , further comprising:
 instructions to process the generated list of one or more subsets of tabular data, wherein the instructions to process include:
 instructions to order the one or more subsets of the generated list, based on the associated calculated priority score; 
 instructions to extract data associated with the one or more subsets of tabular data; 
 instructions to annotate the one or more subsets of tabular data; and 
 instructions to extract data from the annotated one or more subsets of tabular data. 
   
     
     
         18 . The computer program product of  claim 17 , wherein the instructions to extract data associated with the one or more subsets of tabular data occurs subsequently to one or more of: instructions to Extensible Markup Language format; instructions to Unstructured Information Management Architecture format; and instructions to OpenDocument format. 
     
     
         19 . A computer system for generating a plan for document processing, the computer system comprising:
 one or more computer processors;   one or more computer-readable storage media;   program instructions stored on the computer-readable storage media for execution by at least one of the one or more processors, the program instructions comprising:
 instructions to receive a plurality of electronic documents, by a computer using a network; 
 instructions to analyze the received plurality of electronic documents, using the computer, to identify a plurality of tabular data, based on the analyzed plurality of electronic documents; 
 instructions to identify textual data within the identified tabular data, of the analyzed plurality of electronic documents; 
 instructions to generate textual hints, based on the identified textual data within the identified tabular data; 
 instructions to identify references, wherein references are based on matching textual hints with textual data in the received plurality of electronic documents; 
 instructions to calculate a count of references associated with one or more sets of tabular data; 
 instructions to calculate a priority score based on the count of references; and 
 instructions to generate an ingestion plan, based on the calculated priority score. 
   
     
     
         20 . The system of  claim 19 , further comprising:
 instructions to calculate a percent value based on a count of identified textual data associated with the subset one or more subsets of identified tabular data; and   instructions to modify the calculated priority score, based on the calculated percent value.

Join the waitlist — get patent alerts

Track US2017116194A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.