US12135737B1ActiveUtility

Graphical user interface and pipeline for text analytics

Assignee: SAS INST INCPriority: Jun 21, 2023Filed: Mar 25, 2024Granted: Nov 5, 2024
Est. expiryJun 21, 2043(~16.9 yrs left)· nominal 20-yr term from priority
G06F 16/34
69
PatentIndex Score
0
Cited by
74
References
29
Claims

Abstract

A graphical user interface (GUI) and pipeline for processing text documents is provided herein. In one example, a system can receive unstructured text documents. The system can determine entity-issue descriptions corresponding to the unstructured text documents. The system can then generate a GUI indicating the entity-issue descriptions. The GUI can also indicate assignments of the unstructured text documents to categories of a predefined schema. The GUI can allow the user to adjust the assignments of the unstructured text documents to the categories. The GUI can also include a table of rows, where each row corresponds to one of the unstructured text documents. Each row can indicate an entity-issue description in the corresponding unstructured text document and the categories assigned to the unstructured text document. Each row can also include a graphical button that is selectable to allow the user to view the unstructured text document corresponding to the row.

Claims

exact text as granted — not AI-modified
The invention claimed is: 
     
       1. A system comprising:
 one or more processors; and 
 one or more memories including program code that is executable by the one or more processors for causing the one or more processors to perform operations including:
 receiving unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities; 
 determining entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and 
 
 generating a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
 wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and 
 wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model. 
 
 
     
     
       2. The system of  claim 1 , wherein the operations further include:
 applying the rule-based model to the unstructured text documents, wherein the rule-based model is configured to assign the one or more categories to each of the unstructured text documents based on a predefined set of rules, the predefined set of rules indicating which of the one or more categories to assign to the unstructured text documents based on keywords in the unstructured text documents; 
 wherein the graphical user interface is configured to allow the user to approve or disapprove the one or more categories assigned to each of the unstructured text documents by the rule-based model. 
 
     
     
       3. The system of  claim 1 , wherein the operations further include:
 determining a respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents; and 
 configuring the graphical user interface to indicate the respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents. 
 
     
     
       4. The system of  claim 1 , wherein the operations further include:
 generating a first graphical visualization in the graphical user interface indicating how many times each of the entities is mentioned in the unstructured text documents; and 
 generating a second graphical visualization in the graphical user interface indicating how many times each of the issues is mentioned in the unstructured text documents. 
 
     
     
       5. The system of  claim 1 , wherein the graphical user interface further includes:
 a first button for approving all of the assignments; and 
 a second button for disapproving all of the assignments. 
 
     
     
       6. The system of  claim 1 , wherein the graphical user interface further includes:
 a button that is selectable to display a window indicating the one or more categories that are assignable to the unstructured text documents. 
 
     
     
       7. The system of  claim 1 , wherein the operations further include:
 prior to generating the graphical user interface, executing the rule-based model to automatically assign a first set of categories to the unstructured text documents, the first set of categories being different from the entity-issue descriptions; and 
 generating the graphical user interface such that each row of the graphical user interface further indicates:
 the first set of categories assigned by the rule-based model to the unstructured text document corresponding to the row; and 
 a second set of categories assigned by the user to the unstructured text document corresponding to the row. 
 
 
     
     
       8. The system of  claim 7 , wherein each row of the graphical user interface further indicates one or more keywords that are present in the unstructured text document and that were used by the rule-based model to determine the first set of categories to assign to the unstructured text document, wherein the one or more keywords are different from the first set of categories and the second set of categories. 
     
     
       9. The system of  claim 1 , wherein the operations further involve storing change records indicating manual changes to the assignments by the user over a time period. 
     
     
       10. The system of  claim 1 , wherein the graphical user interface further includes a filtering option that is selectable to filter the unstructured text documents based on a selected category, such that only a subset of the unstructured text documents assigned to the selected category are displayed in the graphical user interface. 
     
     
       11. The system of  claim 1 , wherein the operations further involve executing an automated assignment process configured to assign additional unstructured text documents to the one or more categories, the automated assignment process being performed based on the assignments of the unstructured text documents to the one or more categories, the additional unstructured text documents being different from the unstructured text documents. 
     
     
       12. The system of  claim 11 , wherein the automated assignment process involves:
 for each of the additional unstructured text documents:
 identifying an entity-issue description in the additional unstructured text document; 
 determining at least one unstructured text document that has a same entity-issue description as the entity-issue description in the additional unstructured text document; 
 determining at least one category previously assigned to the at least one unstructured text document; and 
 applying the at least one category to the additional unstructured text document. 
 
 
     
     
       13. The system of  claim 1 , wherein operations further involve:
 generating the training data based on the assignments of the unstructured text documents to the one or more categories; and 
 training the machine-learning model based on the training data. 
 
     
     
       14. The system of  claim 1 , wherein operations further involve:
 tuning the rule-based model by adjusting one or more properties of the rule-based model, based on the assignments of the unstructured text documents to the one or more categories. 
 
     
     
       15. A method comprising:
 receiving, by one or more processors, unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities; 
 determining, by the one or more processors, entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and 
 generating, by the one or more processors, a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
 wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and 
 wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model. 
 
 
     
     
       16. The method of  claim 15 , further comprising:
 applying the rule-based model to the unstructured text documents, wherein the rule-based model is configured to assign the one or more categories to each of the unstructured text documents based on a predefined set of rules, the predefined set of rules indicating which of the one or more categories to assign to the unstructured text documents based on keywords in the unstructured text documents; 
 wherein the graphical user interface is configured to allow the user to approve or disapprove the one or more categories assigned to each of the unstructured text documents by the rule-based model. 
 
     
     
       17. The method of  claim 15 , further comprising:
 determining a respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents; and 
 configuring the graphical user interface to indicate the respective frequency at which each of the entity-issue descriptions occurs in the unstructured text documents. 
 
     
     
       18. The method of  claim 15 , further comprising:
 generating a first graphical visualization in the graphical user interface indicating how many times each of the entities is mentioned in the unstructured text documents; and 
 generating a second graphical visualization in the graphical user interface indicating how many times each of the issues is mentioned in the unstructured text documents. 
 
     
     
       19. The method of  claim 15 , wherein the graphical user interface further includes:
 a first button for approving all of the assignments; and 
 a second button for disapproving all of the assignments. 
 
     
     
       20. The method of  claim 15 , wherein the graphical user interface further includes:
 a button that is selectable to display a window indicating the one or more categories that are assignable to the unstructured text documents. 
 
     
     
       21. The method of  claim 15 , wherein each row of the graphical user interface further indicates:
 a first set of categories assigned by the user to the unstructured text document corresponding to the row; and 
 a second set of categories assigned by the rule-based model to the unstructured text document corresponding to the row. 
 
     
     
       22. The method of  claim 21 , wherein each row of the graphical user interface further indicates one or more keywords that are present in the unstructured text document and that were used by the rule-based model to determine the second set of categories to assign to the unstructured text document. 
     
     
       23. The method of  claim 15 , further comprising storing change records indicating manual changes to the assignments by the user over a time period. 
     
     
       24. The method of  claim 15 , further comprising executing an automated assignment process configured to assign additional unstructured text documents to the one or more categories, the automated assignment process being performed based on the assignments of the unstructured text documents to the one or more categories, the additional unstructured text documents being different from the unstructured text documents. 
     
     
       25. The method of  claim 24 , wherein the automated assignment process involves:
 for each of the additional unstructured text documents:
 identifying an entity-issue description in the additional unstructured text document; 
 determining at least one unstructured text document that has a same entity-issue description as the entity-issue description in the additional unstructured text document; 
 determining at least one category previously assigned to the at least one unstructured text document; and 
 applying the at least one category to the additional unstructured text document. 
 
 
     
     
       26. The method of  claim 15 , further comprising:
 generating the training data based on the assignments of the unstructured text documents to the one or more categories; and 
 training the machine-learning model based on the training data. 
 
     
     
       27. The method of  claim 15 , further comprising:
 tuning the rule-based model by adjusting one or more properties of the rule-based model, based on the assignments of the unstructured text documents to the one or more categories. 
 
     
     
       28. A non-transitory computer-readable medium comprising program code that is executable by one or more processors for causing the one or more processors to perform operations including:
 receiving unstructured text documents from one or more sources, wherein the unstructured text documents describe issues with entities; 
 determining entity-issue descriptions corresponding to the unstructured text documents, wherein each of the entity-issue descriptions is a text snippet extracted from a corresponding unstructured text document and wherein the text snippet includes a combination of at least two terms, and wherein the combination of at least two terms includes an entity term, an issue term, and a connector term that links the entity term to the issue term; and 
 generating a graphical user interface indicating the entity-issue descriptions corresponding to the unstructured text documents, the graphical user interface indicating assignments of the unstructured text documents to one or more categories of a predefined schema by a user or a rule-based model, the one or more categories being different from the entity-issue descriptions, the graphical user interface being configured to allow the user to adjust the assignments of the unstructured text documents to the one or more categories based on the entity-issue descriptions in the unstructured text documents;
 wherein the graphical user interface includes a table of rows, each row in the table corresponding to one of the unstructured text documents and indicating a respective entity-issue description in the unstructured text document, each row further indicating the one or more categories of the predefined schema assigned to the unstructured text document, and each row of the graphical user interface further including a graphical button that is selectable to allow the user to selectively view the unstructured text document corresponding to the row; and 
 wherein the assignments of the unstructured text documents to the one or more categories are usable to tune the rule-based model or to generate training data for training a machine-learning model that is different than the rule-based model. 
 
 
     
     
       29. The non-transitory computer-readable medium of  claim 28 , wherein the entity-issue descriptions are determined at least in part by:
 determining parts-of-speech associated with the unstructured text documents; and 
 identifying the entity-issue descriptions based on the parts-of-speech.

Join the waitlist — get patent alerts

Track US12135737B1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.