US2022398857A1PendingUtilityA1
Document analysis architecture
Assignee: AON RISK SERVICES INC OF MARYLANDPriority: Jun 10, 2020Filed: Jun 24, 2022Published: Dec 15, 2022
Est. expiryJun 10, 2040(~13.9 yrs left)· nominal 20-yr term from priority
Inventors:Samuel Cameron FlemingDavid Craig AndrewsJared Dirk SolScott BuzanTimothy SeeganChristopher Ali Mirabzadeh
G06F 18/2413G06F 18/214G06N 20/00G06N 5/022G06V 30/19173G06F 16/35G06V 30/19187G06V 30/416G06F 40/279G06V 30/19147G06F 40/284G06V 30/19133G06V 10/945G06N 3/08G06K 9/627G06V 30/414G06K 9/6256G06N 3/096G06N 3/09
61
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods for generation and use of document analysis architectures are disclosed. A model builder component may be utilized to receiving user input data for labeling a set of documents as in class or out of class. That user input data may be utilized to train one or more classification models, which may then be utilized to predict classification of other documents. Trained models may be incorporated into a model taxonomy for searching and use by other users for document analysis purposes.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A method comprising:
receiving documents that include at least one of patents or patent applications; generating first data representing the documents, the first data distinguishing components of the documents; generating a user interface configured to display:
the components of individual ones of the documents; and
an element configured to accept user input indicating whether the individual ones of the documents are in class or out of class;
generating a classification model based at least in part on user input data corresponding to the user input, the classification model trained utilizing at least a first portion of the documents indicated to be in class by the user input data; and causing the user interface to display an indication of:
the first portion of the documents;
a second portion of the documents marked as out of class in response to the user input;
a third portion of the documents determined to be in class utilizing the classification model; and
a fourth portion of the documents determined to be out of class utilizing the classification model.
3 . The method of claim 2 , wherein the user input comprises first user input, the user input data comprises first user input data, and the method further comprises:
determining a first confidence value associated with results of the classification model; receiving second user input indicating classification of at least one document determined to be in class utilizing the classification model; causing the classification model to be retrained based at least in part on second user input data corresponding to the second user input; determining a second confidence value associated with results of the classification model as retained; and causing display of a trendline representing a change from the first confidence value to the second confidence value.
4 . The method of claim 2 , wherein the user input comprises first user input, the user input data comprises first user input data, and the method further comprises:
receiving second user input indicating classification of at least one document determined to be in class utilizing the classification model; causing the classification model to be retrained based at least in part on second user input data corresponding to the second user input; determining a change in a number of the third portion of the documents marked in class utilizing the classification model as retrained; and causing display of an influence value of the second user input on output by the classification model, the influence value indicating a likelihood that additional user input will impact performance of the classification model.
5 . The method of claim 2 , further comprising:
generating second data indicating a relationship between a first document of the documents and a second document of the documents, the relationship indicating that the first document includes at least one component that is similar to a component of the second document; determining that the user input data indicates that the first document is in class; and determining that the second document is in class based at least in part on the second data indicating the relationship.
6 . A system, comprising:
one or more processors; and non-transitory computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
generating first data representing documents received from one or more databases;
causing display of an element configured to accept user input indicating whether individual ones of the documents are in class or out of class;
generating a model based at least in part on user input data corresponding to the user input;
determining, utilizing the model, a first portion of the documents that are in class and a second portion of the documents that are out of class; and
causing display of an indication of the first portion of the documents and the second portion of the documents.
7 . The system of claim 6 , wherein the user input comprises first user input, the user input data comprises first user input data, and the operations further comprise:
determining a first confidence value associated with results of the model; receiving second user input indicating classification of at least one document determined to be in class utilizing the model; causing the model to be retrained based at least in part second user input data corresponding to the second user input; determining a second confidence value associated with results of the model as retained; and causing display of a trendline representing a change from the first confidence value to the second confidence value.
8 . The system of claim 6 , wherein the user input comprises first user input, the user input data comprises first user input data, and the operations further comprise:
receiving second user input indicating classification of at least one document determined to be in class utilizing the model; causing the model to be retrained based at least in part second user input data corresponding to the second user input; determining a change in a number of at least one of the first portion of the documents or the second portion of the documents utilizing the model as retrained; and causing display of an influence value of the second user input on output by the model.
9 . The system of claim 6 , the operations further comprising:
generating second data indicating a relationship between a first document of the documents and a second document of the documents; determining that the user input data indicates that the first document is in class; and determining that the second document is in class based at least in part on the second data indicating the relationship.
10 . The system of claim 6 , the operations further comprising:
determining, for individual ones of the documents marked as in class, a confidence value; and determining a ranking of the individual ones of the documents marked as in class based at least in part on the confidence value.
11 . The system of claim 6 , the operations further comprising causing display of:
a first indication of a first number of the documents marked in class; a second indication of a second number of the documents marked out of class; a third indication of a third number of the documents determined to be in class utilizing the model; and a fourth indication of a fourth number of the documents determined to be out of class utilizing the model.
12 . The system of claim 6 , the operations further comprising:
causing display of:
a first section indicating first keywords determined to be statistically relevant by the model for identifying the first portion of the documents; and
a second section indicating second keywords determined to be statistically relevant by the model for identifying the second portion of the documents; and
based at least in part on receiving the user input indicating that at least one of the first keywords or the second keywords should be removed, retraining the model to account for removal of the at least one of the first keywords or the second keywords.
13 . The system of claim 6 , the operations further comprising:
searching, utilizing the model, one or more databases for additional documents determined to be in class by the model; receiving an instance of at least one additional documents from the one or more databases; receiving user input indicating classification of the at least one additional documents; and retraining the model based at least in part on the user input indicating the classification of the additional documents.
14 . A method, comprising:
generating first data representing documents received from one or more databases; causing display of an element configured to accept user input indicating whether individual ones of the documents are in class or out of class; generating a model based at least in part on user input data corresponding to the user input; and determining, based at least in part on the model and the user input data, a first portion of the documents that are in class and a second portion of the documents that are out of class.
15 . The method of claim 14 , wherein the user input comprises first user input, the user input data comprises first user input data, and the method further comprises:
determining a first confidence value associated with results of the model; receiving second user input indicating classification of at least one document determined to be in class utilizing the model; causing the model to be retrained based at least in part second user input data corresponding to the second user input; determining a second confidence value associated with results of the model as retained; and causing display of a trendline representing a change from the first confidence value to the second confidence value.
16 . The method of claim 14 , wherein the user input comprises first user input, the user input data comprises first user input data, and the method further comprises:
receiving second user input indicating classification of at least one document determined to be in class utilizing the model; causing the model to be retrained based at least in part second user input data corresponding to the second user input; determining a change in a number of the second portion of the documents determined to be in class utilizing the model as retrained; and causing display of an influence value of the second user input on output by the model.
17 . The method of claim 14 , further comprising:
generating second data indicating a relationship between a first document of the documents and a second document of the documents; determining that the user input data indicates that the first document is in class; and determining that the second document is in class based at least in part on the second data indicating the relationship.
18 . The method of claim 14 , further comprising:
determining, for individual ones of the documents marked as in class, a confidence value indicating a degree of confidence that the individual ones of the documents were marked correctly as in class; and determining a ranking of the individual ones of the documents marked as in class based at least in part on the confidence value.
19 . The method of claim 14 , further comprising causing display of:
a first indication of a first number of the documents marked in class in response to user input; a second indication of a second number of the documents marked out of class in response to the user input; a third indication of a third number of the documents determined to be in class utilizing the model; and a fourth indication of a fourth number of the documents determined to be out of class utilizing the model.
20 . The method of claim 14 , further comprising:
causing display of:
a first section indicating first keywords determined to be statistically relevant by the model for identifying the first portion of the documents; and
a second section indicating second keywords determined to be statistically relevant by the model for identifying the second portion of the documents; and
based at least in part on receiving the user input indicating that at least one of the first keywords or the second keywords should be removed, retraining the model to account for removal of the at least one of the first keywords or the second keywords.
21 . The method of claim 14 , further comprising:
searching, utilizing the model, one or more databases for additional documents determined to be in class by the model; receiving an instance of the additional documents from the one or more databases; receiving user input indicating classification of the additional documents; and retraining the model based at least in part on the user input indicating the classification of the additional documents.Join the waitlist — get patent alerts
Track US2022398857A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.