US2014081995A1PendingUtilityA1

Method and System for Creating a Data Profile Engine, Tool Creation Engines and Product Interfaces for Identifying and Analyzing File and Sections of Files

Assignee: Kiiac LLCPriority: Sep 15, 2008Filed: Nov 6, 2013Published: Mar 20, 2014
Est. expirySep 15, 2028(~2.1 yrs left)· nominal 20-yr term from priority
G06F 16/334G06F 16/355G06F 16/335G06F 17/30699
41
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A data profile engine identifies, classifies, analyzes, searches, compares and cross-references entire files and sections of files, records and other forms of electronic media, and a tool creation engine in combination with the data profile engine builds custom solutions and product interfaces.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for processing documents, comprising:
 generating, by a processor, a word universe comprising statistics of text words in a plurality of documents;   deconstructing each document in the plurality of documents into one or more text blocks;   grouping the text blocks within each document and across the plurality of documents into one or more text block groups based on word characteristics of each text block and match characteristics between the text blocks compared to the word universe; and   generating a data profile for a respective text block group, the data profile comprising word characteristics and match characteristics of the text blocks in the respective text block group.   
     
     
         2 . The method of  claim 1 , wherein the statistics of text words in the word universe comprises word frequencies, word weights, and statistical information including minimum, maximum, mean, median and deviation values. 
     
     
         3 . The method of  claim 1 , wherein deconstructing each document comprises:
 identifying one or more text block demarcations in the document, the demarcations including carriage returns, line feeds, paragraph breaks, headings, captions, prefixes and postfixes; and   dividing the document into one or more text blocks based on the identified text block demarcations.   
     
     
         4 . The method of  claim 1 , wherein the word characteristics of each text block comprise a base score calculated based on a weighted sum of word frequency of each word in the text block compared to the word universe. 
     
     
         5 . The method of  claim 1 , wherein the match characteristics of the text blocks comprise a match score calculated based on a weighted sum of matching words between and among text blocks compared to the word universe. 
     
     
         6 . The method of  claim 1 , further comprising:
 merging one or more text block groups with matching data profiles into a text block group set; and   generating a data profile for the text block group set based on the data profiles of the one or more text block groups.   
     
     
         7 . The method of  claim 6 , further comprising:
 reiterating the document deconstructing, text block grouping, and data profile generating based on the generated data profiles for the one or more text block groups and text block group sets.   
     
     
         8 . The method of  claim 1 , further comprising:
 receiving a source document for analysis;   deconstructing the source document into one or more text blocks;   comparing word characteristics and match characteristics of the one or more text blocks associated with the source document with the generated data profiles for the text block groups to identify matching text block groups; and   determining similarity and divergence between the one or more text blocks associated with the source document and the identified matching text block groups.   
     
     
         9 . The method of  claim 1 , further comprising:
 identifying a set of default clauses from text blocks in a text block group;   identifying one or more alternative clauses for each default clause;   identifying an outline from the text blocks in the text block group; and   generates a template based on the identified outline, default clauses and alternative clauses for the text block group.   
     
     
         10 . The method of  claim 1 , further comprising:
 receiving a search query, the query comprising text terms, clauses and/or text blocks;   comparing word characteristics of the search query with the generated data profiles for the text block groups to identify matching text block groups;   ranking text blocks in the match text block groups based on match characteristics between the search query and the text blocks; and   presenting the ranked text blocks as a search result.   
     
     
         11 . A non-transitory computer-readable storage medium storing executable computer program instructions for processing documents, the computer program instructions comprising instructions for:
 generating, by a processor, a word universe comprising statistics of text words in a plurality of documents;   deconstructing each document in the plurality of documents into one or more text blocks;   grouping the text blocks within each document and across the plurality of documents into one or more text block groups based on word characteristics of each text block and match characteristics between the text blocks compared to the word universe; and   generating a data profile for a respective text block group, the data profile comprising word characteristics and match characteristics of the text blocks in the respective text block group.   
     
     
         12 . The non-transitory computer-readable storage medium of  claim 11 , wherein the statistics of text words in the word universe comprises word frequencies, word weights, and statistical information including minimum, maximum, mean, median and deviation values. 
     
     
         13 . The non-transitory computer-readable storage medium of  claim 11 , wherein the computer program instructions for deconstructing each document further comprises instructions for:
 identifying one or more text block demarcations in the document, the demarcations including carriage returns, line feeds, paragraph breaks, headings, captions, prefixes and postfixes; and   dividing the document into one or more text blocks based on the identified text block demarcations.   
     
     
         14 . The non-transitory computer-readable storage medium of  claim 11 , wherein the word characteristics of each text block comprise a base score calculated based on a weighted sum of word frequency of each word in the text block compared to the word universe. 
     
     
         15 . The non-transitory computer-readable storage medium of  claim 11 , wherein the match characteristics of the text blocks comprise a match score calculated based on a weighted sum of matching words between and among text blocks compared to the word universe. 
     
     
         16 . The non-transitory computer-readable storage medium of  claim 11 , wherein the computer program instructions further comprise instructions for:
 merging one or more text block groups with matching data profiles into a text block group set; and   generating a data profile for the text block group set based on the data profiles of the one or more text block groups.   
     
     
         17 . The non-transitory computer-readable storage medium of  claim 16 , wherein the computer program instructions further comprise instructions for:
 reiterating the document deconstructing, text block grouping, and data profile generating based on the generated data profiles for the one or more text block groups and text block group sets.   
     
     
         18 . The non-transitory computer-readable storage medium of  claim 11 , wherein the computer program instructions further comprise instructions for:
 receiving a source document for analysis;   deconstructing the source document into one or more text blocks;   comparing word characteristics and match characteristics of the one or more text blocks associated with the source document with the generated data profiles for the text block groups to identify matching text block groups; and   determining similarity and divergence between the one or more text blocks associated with the source document and the identified matching text block groups.   
     
     
         19 . The non-transitory computer-readable storage medium of  claim 11 , wherein the computer program instructions further comprise instructions for:
 identifying a set of default clauses from text blocks in a text block group;   identifying one or more alternative clauses for each default clause;   identifying an outline from the text blocks in the text block group; and   generates a template based on the identified outline, default clauses and alternative clauses for the text block group.   
     
     
         20 . The non-transitory computer-readable storage medium of  claim 11 , wherein the computer program instructions further comprise instructions for:
 receiving a search query, the query comprising text terms, clauses and/or text blocks;   comparing word characteristics of the search query with the generated data profiles for the text block groups to identify matching text block groups;   ranking text blocks in the match text block groups based on match characteristics between the search query and the text blocks; and   presenting the ranked text blocks as a search result.

Join the waitlist — get patent alerts

Track US2014081995A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.