US2022036244A1PendingUtilityA1

Systems and methods for predictive coding

Assignee: OPEN TEXT HOLDINGS INCPriority: May 25, 2010Filed: Oct 18, 2021Published: Feb 3, 2022
Est. expiryMay 25, 2030(~3.8 yrs left)· nominal 20-yr term from priority
G06N 7/01G06F 16/93G06N 5/048G06N 20/10G06N 5/04G06N 20/00G06N 7/005
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 receiving a hard coding correction to at least a portion of an automated coding of a relevant document;   updating a coded control set with the relevant document having the hard coding correction;   applying the updated coded control set to another set of additional documents to automatically code the another set of additional documents;   identifying contextually similar documents to an initial set of relevant documents utilizing machine learning to automatically detect concepts within the initial set of relevant documents using a statistical analysis; and   supplementing the initial set of relevant documents with the contextually similar documents.   
     
     
         2 . The method according to  claim 1 , further comprising receiving training data that comprises an initial set of relevant documents including computer-suggested documents from machine learning. 
     
     
         3 . The method according to  claim 1 , further comprising generating a coded control set of data based on training data, using coding determinations for the coded control set, the coding determinations comprising automated coding of the relevant documents performed by a predictive coding system. 
     
     
         4 . The method according to  claim 1 , further comprising automatically coding the another set of additional documents with the coded control set using a predictive coding system. 
     
     
         5 . The method according to  claim 1 , further comprising presenting a relevant document from a set of additional relevant documents to a human reviewer, the relevant document having the automated coding from a predictive coding system. 
     
     
         6 . The method according to  claim 1 , further comprising allowing a human reviewer to correct at least a portion of the automated coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the automated coding from a first coding to a second coding; 
     
     
         7 . The method according to  claim 1 , further comprising applying coding determinations of the coded control set to contextually similar data in a corpus of documents. 
     
     
         8 . The method according to  claim 1 , further comprising receiving a review of each document in the updated coded control set to ensure that proper coding has been implemented. 
     
     
         9 . The method according to  claim 1 , further comprising:
 generating a coded control set of data based on a plurality of uncoded documents from a corpus, the plurality of uncoded documents comprising documents that are relevant or selected randomly, wherein the portion of the documents are randomly sampled from an un-reviewed document population; and   receiving a determination if the document coded using the coded control set was miscoded, prior to the step of receiving the hard coding correction to the document.   
     
     
         10 . The method of  claim 1 , wherein updating the coded control set comprises an adaptive identification life cycle that receives the hard coding correction and updates the coded control set with the relevant document, wherein the adaptive identification life cycle terminates based on a confidence threshold validation applied to the updated automatically coded control set and the automatically coded another set of additional documents. 
     
     
         11 . A system comprising:
 a processor and memory for storing instructions, the processor executing the instructions to:
 receive a hard coding correction to at least a portion of an automated coding of a relevant document, the hard coding correction being produced by a human reviewer; 
 update a coded control set with the relevant document having the hard coding correction; 
 apply the updated coded control set to another set of additional documents to automatically code the additional documents; 
 identify contextually similar documents to an initial set of relevant documents utilizing machine learning to automatically detect concepts within the initial set of relevant documents using a statistical analysis; and 
 supplement the initial set of relevant documents with the contextually similar documents. 
   
     
     
         12 . The system according to  claim 11 , wherein the processor is configured to receive training data that comprises an initial set of relevant documents including computer-suggested documents from machine learning. 
     
     
         13 . The system according to  claim 11 , wherein the processor is configured to generate coded control set based on training data using coding determinations for the coded control set, the coding determinations comprising automated coding of the relevant documents performed by a predictive coding system. 
     
     
         14 . The system according to  claim 11 , wherein the processor is configured to automaticaly code the another set of additional documents with the coded control set using a predictive coding system. 
     
     
         15 . The system according to  claim 11 , wherein the processor is configured to present a relevant document from another set of additional relevant documents to a human reviewer, the relevant document having the automated coding from a predictive coding system. 
     
     
         16 . The system according to  claim 11 , wherein the processor is configured to allow the human reviewer to correct at least a portion of the automated coding by performing a hard coding correction, the hard coding correction comprising changing of the at least a portion of the automated coding from a first coding to a second coding; 
     
     
         17 . The system according to  claim 11 , wherein the processor is configured to apply coding determinations of the coded control set to contextually similar data in a corpus of documents. 
     
     
         18 . The system according to  claim 11 , wherein the processor is configured to receive a review of each document in the updated coded control set to ensure that proper coding has been implemented. 
     
     
         19 . The system according to  claim 18 , wherein the processor generates the coded control set based on a plurality of uncoded documents from a corpus, the plurality of uncoded documents comprising documents that are relevant or selected randomly, wherein the portion of the documents are randomly sampled from an un-reviewed document population. 
     
     
         20 . The system according to  claim 11 , wherein the processor is configured to receive a determination if the document coded using the coded control set was miscoded, prior to the step of receiving the hard coding correction to the document.

Join the waitlist — get patent alerts

Track US2022036244A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.