US2014207786A1PendingUtilityA1

System and methods for computerized information governance of electronic documents

Assignee: EQUIVIO LTDPriority: Jan 22, 2013Filed: Oct 24, 2013Published: Jul 24, 2014
Est. expiryJan 22, 2033(~6.5 yrs left)· nominal 20-yr term from priority
G06F 16/285G06F 16/335G06F 16/288G06F 17/30598
44
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An information governance system comprising a plurality of classifiers which employ cutoffs for classifying at least a portion of a population of incoming documents as documents to be retained and documents to be discarded in accordance with a corresponding plurality of pre-defined retention schedules; training apparatus for training said classifiers based on relevance inputs provided by a human information governance expert regarding a training set of documents within a universe of documents to be governed; and apparatus operative to automatically cause any classified document to be retained and subsequently discarded in accordance with its pre-defined retention schedule including discarding only documents that (a) have been classified as documents to be discarded and (b) have not been classified as documents to be retained, and to automatically cause any document which could not be classified, to be retained as gray area data until further notice.

Claims

exact text as granted — not AI-modified
1 . An information governance system comprising:
 A plurality of classifiers which employ cutoffs for classifying at least a portion of a population of incoming documents as documents to be retained and documents to be discarded in accordance with a retention policy comprising a corresponding plurality of pre-defined retention schedules;   training apparatus for training said classifiers based on relevance inputs provided by a human information governance expert regarding a training set of documents within a universe of documents to be governed; and   retain/discard apparatus operative to automatically cause any classified document to be retained and subsequently discarded in accordance with its pre-defined retention schedule including discarding only documents that (a) have been classified as documents to be discarded and (b) have not been classified as documents to be retained, and to automatically cause any document which could not be classified, to be retained as gray area data until further notice.   
     
     
         2 . A system according to  claim 1  and also comprising computerized apparatus for identifying at least one cluster of related documents within a set of “gray area” documents which could not be classified. 
     
     
         3 . A system according to  claim 1  wherein at least one of the plurality of pre-defined retention schedules calls for documents to be discarded immediately. 
     
     
         4 . A system according to  claim 1  wherein the training apparatus is operative to train each of said classifiers until a predetermined precision measure has been achieved. 
     
     
         5 . A system according to  claim 1  and also comprising threshold adjustment functionality operative to quantify a false-negative error rate, resulting in premature discarding of documents, and, if excessive, to adjust at least one threshold employed by said classifiers accordingly. 
     
     
         6 . A system according to  claim 5  wherein a false-negative error rate is deemed excessive if a pre-stored human categorizer's false-negative error rate is lower. 
     
     
         7 . A system according to  claim 1  and also comprising identifying and discarding older near-duplicates of at least one retained document. 
     
     
         8 . An information governance method comprising:
 generating a plurality of classifiers for classifying electronic documents into a corresponding plurality of documentation retention categories;   running training iterations thereby to improve at least one of the plurality of classifiers;   classifying a repository of electronic documents using said plurality of classifiers and running a Logarithmic stratified sampling-based Quality Assurance process to compute precision in cases of low or unknown richness including ordering documents by their ranks then partitioning the ranks into slices: [0,p] [p, 2p], [2p, 4p], . . . , and randomly selecting documents to represent each slice, thereby to generate Quality Assurance results;   if the Quality Assurance results are not deemed good enough, improve the classifier and return to one of said running steps:   if the Quality Assurance results are good enough, use last classifier to implement a plurality of document retention settings corresponding to said plurality of documentation retention categories.   
     
     
         9 . A method according to  claim 8  and also comprising:
 using a processor for identifying at least one cluster of related documents within a set of “gray area” documents which could not be classified; 
 generating at least one additional classifier for classifying electronic documents into each of at least one cluster of related documents;
 repeating said classifying step using both the plurality of classifiers and the at least one additional classifier thereby to reduce percentage of “gray area” documents; and 
 implementing document retention settings including:
 settings corresponding to said plurality of documentation retention categories and 
 settings corresponding to each of said at least one cluster of related documents. 
 
 
 
     
     
         10 . A method according to  claim 8  wherein for each individual classifier, stability measures are computed based on cross validation and percentage of relevant documents in the individual classifier's training set. 
     
     
         11 . A method according to  claim 9  wherein said using a processor for identifying comprises using Equivio Themes functionality for identifying at least one cluster of related documents within a set of “gray area” documents which could not be classified. 
     
     
         12 . A system according to  claim 1  and also comprising a rule repository operative to map said retention policy to retention time and wherein said rules are accessed by said retain/discard apparatus. 
     
     
         13 . A method according to  claim 8  wherein if the Quality Assurance results are not deemed good enough for an individual classifier, the individual classifier is improved by adding documents used in said quality assurance process to the individual classifier's training set thereby to generate an expanded training set and re-training the individual classifier using the expanded training set. 
     
     
         14 . A method according to  claim 8  wherein said plurality of documentation retention categories includes at least one category of documents to be retained and at least one category of documents to be immediately discarded. 
     
     
         15 . A method according to  claim 14  wherein for each classifier from among said plurality of classifiers which corresponds to a category of documents to be retained high and low cutoff points are set. 
     
     
         16 . A method according to  claim 15  wherein for each classifier from among said plurality of classifiers which corresponds to a category of documents to be immediately discarded, just one cutoff is set. 
     
     
         17 . A method according to  claim 15  wherein an individual document is:
 retained if the individual document's relevance exceeds the high cutoff of a retention category to which said individual document belongs, and 
 discarded if the document's relevance both falls above a cutoff of a category of documents to be immediately discarded to which the individual document belongs and falls below all low cutoff points of all retention categories. 
 
     
     
         18 . A method according to  claim 17  wherein the individual document is retained as a gray area document if the individual document is not discarded and if the individual document falls below the high cutoff of the retention category to which said individual document belongs. 
     
     
         19 . A computer program product, comprising a non-transitory tangible computer readable medium having computer readable program code embodied therein, said computer readable program code adapted to be executed to implement an information governance method comprising:
 generating a plurality of classifiers for classifying electronic documents into a corresponding plurality of documentation retention categories;   running training iterations thereby to improve at least one of the plurality of classifiers;   classifying a repository of electronic documents using said plurality of classifiers and running a Logarithmic stratified sampling-based Quality Assurance process to compute precision in cases of low or unknown richness including ordering documents by their ranks then partitioning the ranks into slices: [0,p] [p, 2p], [2p, 4p], . . . , and randomly selecting documents to represent each slice, thereby to generate Quality Assurance results;   if the Quality Assurance results are not deemed good enough, improve the classifier and return to one of said running steps;   if the Quality Assurance results are good enough, use last classifier to implement a plurality of document retention settings corresponding to said plurality of documentation retention categories.

Join the waitlist — get patent alerts

Track US2014207786A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.