Graphical user interface and pipeline for text analytics
Abstract
A graphical user interface (GUI) and pipeline for processing text documents is provided herein. In one example, a system can receive unstructured text documents. The system can determine entity-issue descriptions corresponding to the unstructured text documents. The system can then generate a GUI indicating the entity-issue descriptions. The GUI can also indicate assignments of the unstructured text documents to categories of a predefined schema. The GUI can allow the user to adjust the assignments of the unstructured text documents to the categories. The GUI can also include a table of rows, where each row corresponds to one of the unstructured text documents. Each row can indicate an entity-issue description in the corresponding unstructured text document and the categories assigned to the unstructured text document. Each row can also include a graphical button that is selectable to allow the user to view the unstructured text document corresponding to the row.
Claims
exact text as granted — not AI-modifiedThe invention claimed is:
1. A system comprising:
one or more processors; and
one or more memories including program code that is executable by the one or more processors for causing the one or more processors to perform operations including:
receiving unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities;
determining entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and
generating a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and
wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model.
2. The system of claim 1 , wherein the operations further include:
applying the rule-based model to the unstructured text documents, wherein the rule-based model is configured to assign the one or more categories to each of the unstructured text documents based on a predefined set of rules, the predefined set of rules indicating which of the one or more categories to assign to the unstructured text documents based on keywords in the unstructured text documents;
wherein the graphical user interface is configured to allow the user to approve or disapprove the one or more categories assigned to each of the unstructured text documents by the rule-based model.
3. The system of claim 1 , wherein the operations further include:
determining a respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents; and
configuring the graphical user interface to indicate the respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents.
4. The system of claim 1 , wherein the operations further include:
generating a first graphical visualization in the graphical user interface indicating how many times each of the entities is mentioned in the unstructured text documents; and
generating a second graphical visualization in the graphical user interface indicating how many times each of the issues is mentioned in the unstructured text documents.
5. The system of claim 1 , wherein the graphical user interface further includes:
a first button for approving all of the assignments; and
a second button for disapproving all of the assignments.
6. The system of claim 1 , wherein the graphical user interface further includes:
a button that is selectable to display a window indicating the one or more categories that are assignable to the unstructured text documents.
7. The system of claim 1 , wherein the operations further include:
prior to generating the graphical user interface, executing the rule-based model to automatically assign a first set of categories to the unstructured text documents, the first set of categories being different from the entity-issue descriptions; and
generating the graphical user interface such that each row of the graphical user interface further indicates:
the first set of categories assigned by the rule-based model to the unstructured text document corresponding to the row; and
a second set of categories assigned by the user to the unstructured text document corresponding to the row.
8. The system of claim 7 , wherein each row of the graphical user interface further indicates one or more keywords that are present in the unstructured text document and that were used by the rule-based model to determine the first set of categories to assign to the unstructured text document, wherein the one or more keywords are different from the first set of categories and the second set of categories.
9. The system of claim 1 , wherein the operations further involve storing change records indicating manual changes to the assignments by the user over a time period.
10. The system of claim 1 , wherein the graphical user interface further includes a filtering option that is selectable to filter the unstructured text documents based on a selected category, such that only a subset of the unstructured text documents assigned to the selected category are displayed in the graphical user interface.
11. The system of claim 1 , wherein the operations further involve executing an automated assignment process configured to assign additional unstructured text documents to the one or more categories, the automated assignment process being performed based on the assignments of the unstructured text documents to the one or more categories, the additional unstructured text documents being different from the unstructured text documents.
12. The system of claim 11 , wherein the automated assignment process involves:
for each of the additional unstructured text documents:
identifying an entity-issue description in the additional unstructured text document;
determining at least one unstructured text document that has a same entity-issue description as the entity-issue description in the additional unstructured text document;
determining at least one category previously assigned to the at least one unstructured text document; and
applying the at least one category to the additional unstructured text document.
13. The system of claim 1 , wherein operations further involve:
generating the training data based on the assignments of the unstructured text documents to the one or more categories; and
training the machine-learning model based on the training data.
14. The system of claim 1 , wherein operations further involve:
tuning the rule-based model by adjusting one or more properties of the rule-based model, based on the assignments of the unstructured text documents to the one or more categories.
15. A method comprising:
receiving, by one or more processors, unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities;
determining, by the one or more processors, entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and
generating, by the one or more processors, a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and
wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model.
16. The method of claim 15 , further comprising:
applying the rule-based model to the unstructured text documents, wherein the rule-based model is configured to assign the one or more categories to each of the unstructured text documents based on a predefined set of rules, the predefined set of rules indicating which of the one or more categories to assign to the unstructured text documents based on keywords in the unstructured text documents;
wherein the graphical user interface is configured to allow the user to approve or disapprove the one or more categories assigned to each of the unstructured text documents by the rule-based model.
17. The method of claim 15 , further comprising:
determining a respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents; and
configuring the graphical user interface to indicate the respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents.
18. The method of claim 15 , further comprising:
generating a first graphical visualization in the graphical user interface indicating how many times each of the entities is mentioned in the unstructured text documents; and
generating a second graphical visualization in the graphical user interface indicating how many times each of the issues is mentioned in the unstructured text documents.
19. The method of claim 15 , wherein the graphical user interface further includes:
a first button for approving all of the assignments; and
a second button for disapproving all of the assignments.
20. The method of claim 15 , wherein the graphical user interface further includes:
a button that is selectable to display a window indicating the one or more categories that are assignable to the unstructured text documents.
21. The method of claim 15 , wherein each row of the graphical user interface further indicates:
a first set of categories assigned by the user to the unstructured text document corresponding to the row; and
a second set of categories assigned by the rule-based model to the unstructured text document corresponding to the row.
22. The method of claim 21 , wherein each row of the graphical user interface further indicates one or more keywords that are present in the unstructured text document and that were used by the rule-based model to determine the second set of categories to assign to the unstructured text document.
23. The method of claim 15 , further comprising storing change records indicating manual changes to the assignments by the user over a time period.
24. The method of claim 15 , further comprising executing an automated assignment process configured to assign additional unstructured text documents to the one or more categories, the automated assignment process being performed based on the assignments of the unstructured text documents to the one or more categories, the additional unstructured text documents being different from the unstructured text documents.
25. The method of claim 24 , wherein the automated assignment process involves:
for each of the additional unstructured text documents:
identifying an entity-issue description in the additional unstructured text document;
determining at least one unstructured text document that has a same entity-issue description as the entity-issue description in the additional unstructured text document;
determining at least one category previously assigned to the at least one unstructured text document; and
applying the at least one category to the additional unstructured text document.
26. The method of claim 15 , further comprising:
generating the training data based on the assignments of the unstructured text documents to the one or more categories; and
training the machine-learning model based on the training data.
27. The method of claim 15 , further comprising:
tuning the rule-based model by adjusting one or more properties of the rule-based model, based on the assignments of the unstructured text documents to the one or more categories.
28. A non-transitory computer-readable medium comprising program code that is executable by one or more processors for causing the one or more processors to perform operations including:
receiving unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities;
determining entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and
generating a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and
wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model.
29. The non-transitory computer-readable medium of claim 28 , wherein the entity-issue descriptions are determined at least in part by:
determining parts-of-speech associated with the unstructured text documents; and
identifying the entity-issue descriptions based on the parts-of-speech.Join the waitlist — get patent alerts
Track US12135737B1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.