Continuous learning, prediction, and ranking of relevancy or non-relevancy of discovery documents using a caseassist active learning and dynamic document review workflow
Abstract
A method includes receiving a set of documents associated with data discovery. The method further includes receiving, for each document in a subset of the set of documents, an indication of relevancy or non-relevancy of the document for an issue. The method further includes modifying one or more parameters for a machine-learning model based on the indication of relevancy or non-relevancy. The method further includes outputting, for each document in the set of documents, by a machine-learning model, a prediction probability of relevancy to an issue associated with the data discovery, and a ranking of the set of documents based on the prediction probability of relevancy. The method further includes generating a user interface that includes a sampling of the documents for review by a user, where each document is associated with a predicted relevancy tag or a predicted non-relevancy tag.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method to train a machine-learning model, the method comprising:
receiving a set of training documents associated with data discovery; outputting for each document in the set of training documents, by the machine-learning model, a prediction probability of relevancy to an issue associated with the data discovery, and a ranking of the set of training documents based on the prediction probability of relevancy; generating a user interface that includes a sampling of the documents for review by a user, wherein each document in the sampling of the documents is associated with a predicted relevancy tag or a predicted non-relevancy tag; receiving user input that modifies the predicted relevancy tag or the predicted non-relevancy tag for one or more documents in the sampling of documents; modifying, as a first iteration, the machine-learning model based on the user input; and in response to modifying the machine-learning model, generating a reranking of the set of training documents.
2 . The method of claim 1 , further comprising:
updating the user interface to include an estimated document review time of each iteration, wherein the estimated document review time includes at least one selected from the group of a number of documents associated with the predicted relevancy tag, a number of documents reviewed per hour, a number of hours worked per day, a number of days to complete the review, and combinations thereof.
3 . The method of claim 1 , further comprising:
updating the user interface to include a total number of documents reviewed by the user, a number of documents reviewed relevant as compared to a number of documents predicted as being non-relevant from the number of documents reviewed relevant, and a number of documents reviewed non-relevant as compared to a number of documents predicted as being relevant.
4 . The method of claim 1 , further comprising:
updating the user interface to include a graph of a number of documents that are associated with the predicted relevancy tag or the predicted non-relevancy tag as a function of a prediction probability.
5 . The method of claim 1 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user.
6 . The method of claim 5 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user and an accuracy of reviewed documents from the sampling of the documents falls below a predetermined threshold.
7 . The method of claim 1 , wherein the user interface includes a sampling criteria that includes options for configuring how the machine-learning model is verified, wherein the options include a confidence level associated with the predicted relevancy tag or the predicted non-relevancy tag associated with each document in the sampling and a margin of error.
8 . The method of claim 1 , further comprising:
receiving, as input to a trained machine-learning model, a set of new documents associated with the data discovery; and outputting for each document in the set of new documents, by the trained machine-learning model, the predicted relevancy tag or the predicted non-relevancy tag.
9 . A system to train a machine-learning model, the system comprising:
one or more processors; and a memory coupled to the processor, with instructions stored thereon that, when executed by the processor, cause the processor to perform operations comprising:
receiving a set of training documents associated with data discovery;
outputting for each document in the set of training documents, by the machine-learning model, a prediction probability of relevancy to an issue associated with the data discovery, and a ranking of the set of training documents based on the prediction probability of relevancy;
generating a user interface that includes a sampling of the documents for review by a user, wherein each document in the sampling of the documents is associated with a predicted relevancy tag or a predicted non-relevancy tag;
receiving user input that modifies the predicted relevancy tag or the predicted non-relevancy tag for one or more documents in the sampling of documents;
modifying, as a first iteration, the machine-learning model based on the user input; and
in response to modifying the machine-learning model, generating a reranking of the set of training documents.
10 . The system of claim 9 , wherein the operations further comprise:
updating the user interface to include an estimated document review time of each iteration, wherein the estimated document review time includes at least one selected from the group of a number of documents associated with the predicted relevancy tag, a number of documents reviewed per hour, a number of hours worked per day, a number of days to complete the review, and combinations thereof.
11 . The system of claim 9 , wherein the operations further comprise:
updating the user interface to include a total number of documents reviewed by the user, a number of documents reviewed relevant as compared to a number of documents predicted as being non-relevant from the number of documents reviewed relevant, and a number of documents reviewed non-relevant as compared to a number of documents predicted as being relevant.
12 . The system of claim 9 , wherein the operations further comprise:
updating the user interface to include a graph of a number of documents that are associated with the predicted relevancy tag or the predicted non-relevancy tag as a function of a prediction probability.
13 . The system of claim 9 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user.
14 . The system of claim 13 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user and an accuracy of reviewed documents from the sampling of the documents falls below a predetermined threshold.
15 . A non-transitory computer-readable medium to train a machine-learning model with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:
receiving a set of training documents associated with data discovery; outputting for each document in the set of training documents, by the machine-learning model, a prediction probability of relevancy to an issue associated with the data discovery, and a ranking of the set of training documents based on the prediction probability of relevancy; generating a user interface that includes a sampling of the documents for review by a user, wherein each document in the sampling of the documents is associated with a predicted relevancy tag or a predicted non-relevancy tag; receiving user input that modifies the predicted relevancy tag or the predicted non-relevancy tag for one or more documents in the sampling of documents; modifying, as a first iteration, the machine-learning model based on the user input; and in response to modifying the machine-learning model, generating a reranking of the set of training documents.
16 . The computer-readable medium of 15 , wherein the operations further comprise:
updating the user interface to include an estimated document review time of each iteration, wherein the estimated document review time includes at least one selected from the group of a number of documents associated with the predicted relevancy tag, a number of documents reviewed per hour, a number of hours worked per day, a number of days to complete the review, and combinations thereof.
17 . The computer-readable medium of 15 , wherein the operations further comprise:
updating the user interface to include a total number of documents reviewed by the user, a number of documents reviewed relevant as compared to a number of documents predicted as being non-relevant from the number of documents reviewed relevant, and a number of documents reviewed non-relevant as compared to a number of documents predicted as being relevant.
18 . The computer-readable medium of 15 , wherein the operations further comprise:
updating the user interface to include a graph of a number of documents that are associated with the predicted relevancy tag or the predicted non-relevancy tag as a function of a prediction probability.
19 . The computer-readable medium of 15 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user.
20 . The computer-readable medium of 15 , wherein the user interface includes an option that causes the machine-learning model to be modified after a predetermined number of the sampling of the documents are reviewed by the user and an accuracy of reviewed documents from the sampling of the documents falls below a predetermined threshold.Join the waitlist — get patent alerts
Track US2023102591A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.