US2022318681A1PendingUtilityA1

System and method for scalable, interactive, collaborative topic identification and tracking

Assignee: CAPITAL ONE SERVICES LLCPriority: May 15, 2019Filed: Jun 16, 2022Published: Oct 6, 2022
Est. expiryMay 15, 2039(~12.8 yrs left)· nominal 20-yr term from priority
G06N 7/01G06F 16/904G06F 16/358G06N 20/00
61
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A topic tracking platform is disclosed that includes a machine-learning model that may be trained to expose topics in a corpus in response to a training table. Because topics are exposed, rather than searched for using existing taxonomies, the sensitivity of a topic tracking platform may be increased, and emerging topic trends may be more quickly flagged. Exposed topics may be automatically labelled, increasing the specificity of the topic tracking platform by overcoming the potential for topic labelling inconsistencies currently experienced in the art. Documents may be scored for each topic using information provided at a token granularity, and the contribution that each token of each document contributes to the topic may be visually represented. In some aspects, mechanisms are provided for reviewing topics of the corpus at varying granularities, including at a topic level, document level or token level granularity.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method, comprising:
 determining a corpus of documents and a model to apply the corpus;   iteratively applying the corpus to the model to identify topics in the corpus of documents, each topic associated with one or more words;   determining, for each topic, contribution values for the one or more words to the topic, wherein each contribution value corresponds to one of the one or more words, and each contribution value indicates a frequency one of the one or more words contributes to the topic;   generating a first map comprising the one or more words, the contribution values, and the topics, wherein each word of the one or more words is mapped to a particular topic and has a corresponding contribution value; and   storing, in storage, the first map.   
     
     
         3 . The method of  claim 2 , comprising displaying, the first map in a table format illustrating the one or more words and their contribution values to each topic. 
     
     
         4 . The method of  claim 2 , comprising:
 generating a second map comprising, for each of the documents of the corpus, a relative distribution of the one or more words per topic per document; and   storing, in the storage, the second map.   
     
     
         5 . The method of  claim 4 , comprising determining the relative distribution by summing the contribution value of each word of the document for each topic. 
     
     
         6 . The method of  claim 4 , comprising displaying, the second map in a table format. 
     
     
         7 . The method of  claim 2 , wherein the model is selected from a plurality of models stored in the storage, and the plurality of models are trained using a plurality of different corpora. 
     
     
         8 . The method of  claim 7 , wherein the plurality of different corpora comprise corpora captured from different time periods, different sources, or a combination thereof. 
     
     
         9 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a processor, cause the processor to:
 determine a corpus of documents and a model to apply the corpus;   iteratively apply the corpus to the model to identify topics in the corpus of documents, each topic associated with one or more words;   determine, for each topic, contribution values for the one or more words to the topic, wherein each contribution value corresponds to one of the one or more words, and each contribution value indicates a frequency one of the one or more words contributes to the topic;   generate a first map comprising the one or more words, the contribution values, and the topics, wherein each word of the one or more words is mapped to a particular topic and has a corresponding contribution value; and   store, in the storage, the first map.   
     
     
         10 . The computer-readable storage medium of  claim 9 , comprising display, the first map in a table format illustrating the one or more words and their contribution values to each topic. 
     
     
         11 . The computer-readable storage medium of  claim 9 , comprising:
 generate a second map comprising, for each of the documents of the corpus, a relative distribution of the one or more words per topic per document; and   store, in the storage, the second map.   
     
     
         12 . The computer-readable storage medium of  claim 11 , comprising determine the relative distribution by summing the contribution value of each word of the document for each topic. 
     
     
         13 . The computer-readable storage medium of  claim 11 , comprising display, the second map in a table format. 
     
     
         14 . The computer-readable storage medium of  claim 9 , wherein the model is selected from a plurality of models stored in the storage, and the plurality of models are trained using a plurality of different corpora. 
     
     
         15 . The computer-readable storage medium of  claim 14 , wherein the plurality of different corpora comprise corpora captured from different time periods, different sources, or a combination thereof. 
     
     
         16 . A system comprising:
 storage;   processing circuitry configured to execute instructions, that when executed, cause the processing circuitry to:   determine a corpus of documents and a model to apply the corpus;   iteratively apply the corpus to the model to identify topics in the corpus of documents, each topic comprising a plurality of hierarchically organized components;   determine, for each topic, contribution values for the plurality of hierarchically organized components to the topic, wherein each contribution value corresponds to one of the hierarchically organized components, and each contribution value indicates a frequency one of the plurality of hierarchically organized components contributes to the topic;   determine a first map comprising the plurality of hierarchically organized components, the contribution values, and the topics, wherein each hierarchically organized component is mapped to a particular topic and has a corresponding contribution value; and   store, in the storage, the first map.   
     
     
         17 . The system of  claim 16 , comprising the processing circuitry configured to display, the first map in a table format illustrating the plurality of hierarchically organized components and their contribution values to each topic. 
     
     
         18 . The system of  claim 16 , comprising the processing circuitry configured to:
 generate a second map comprising, for each of the documents of the corpus, a relative distribution of the plurality of hierarchically organized components per topic per document; and   store, in the storage, the second map.   
     
     
         19 . The system of  claim 18 , comprising determine the relative distribution by summing the contribution value of each hierarchically organized components of the document for each topic. 
     
     
         20 . The system of  claim 18 , comprising display, the second map in a table format. 
     
     
         21 . The system of  claim 16 , wherein the plurality of hierarchically organized components comprising words.

Join the waitlist — get patent alerts

Track US2022318681A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.