System and method for scalable, interactive, collaborative topic identification and tracking
Abstract
A topic tracking platform is disclosed that includes a machine-learning model that may be trained to expose topics in a corpus in response to a training table. Because topics are exposed, rather than searched for using existing taxonomies, the sensitivity of a topic tracking platform may be increased, and emerging topic trends may be more quickly flagged. Exposed topics may be automatically labelled, increasing the specificity of the topic tracking platform by overcoming the potential for topic labelling inconsistencies currently experienced in the art. Documents may be scored for each topic using information provided at a token granularity, and the contribution that each token of each document contributes to the topic may be visually represented. In some aspects, mechanisms are provided for reviewing topics of the corpus at varying granularities, including at a topic level, document level or token level granularity.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A method, comprising:
determining a corpus of documents and a model to apply the corpus; iteratively applying the corpus to the model to identify topics in the corpus of documents, each topic associated with one or more words; determining, for each topic, contribution values for the one or more words to the topic, wherein each contribution value corresponds to one of the one or more words, and each contribution value indicates a frequency one of the one or more words contributes to the topic; generating a first map comprising the one or more words, the contribution values, and the topics, wherein each word of the one or more words is mapped to a particular topic and has a corresponding contribution value; and storing, in storage, the first map.
3 . The method of claim 2 , comprising displaying, the first map in a table format illustrating the one or more words and their contribution values to each topic.
4 . The method of claim 2 , comprising:
generating a second map comprising, for each of the documents of the corpus, a relative distribution of the one or more words per topic per document; and storing, in the storage, the second map.
5 . The method of claim 4 , comprising determining the relative distribution by summing the contribution value of each word of the document for each topic.
6 . The method of claim 4 , comprising displaying, the second map in a table format.
7 . The method of claim 2 , wherein the model is selected from a plurality of models stored in the storage, and the plurality of models are trained using a plurality of different corpora.
8 . The method of claim 7 , wherein the plurality of different corpora comprise corpora captured from different time periods, different sources, or a combination thereof.
9 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a processor, cause the processor to:
determine a corpus of documents and a model to apply the corpus; iteratively apply the corpus to the model to identify topics in the corpus of documents, each topic associated with one or more words; determine, for each topic, contribution values for the one or more words to the topic, wherein each contribution value corresponds to one of the one or more words, and each contribution value indicates a frequency one of the one or more words contributes to the topic; generate a first map comprising the one or more words, the contribution values, and the topics, wherein each word of the one or more words is mapped to a particular topic and has a corresponding contribution value; and store, in the storage, the first map.
10 . The computer-readable storage medium of claim 9 , comprising display, the first map in a table format illustrating the one or more words and their contribution values to each topic.
11 . The computer-readable storage medium of claim 9 , comprising:
generate a second map comprising, for each of the documents of the corpus, a relative distribution of the one or more words per topic per document; and store, in the storage, the second map.
12 . The computer-readable storage medium of claim 11 , comprising determine the relative distribution by summing the contribution value of each word of the document for each topic.
13 . The computer-readable storage medium of claim 11 , comprising display, the second map in a table format.
14 . The computer-readable storage medium of claim 9 , wherein the model is selected from a plurality of models stored in the storage, and the plurality of models are trained using a plurality of different corpora.
15 . The computer-readable storage medium of claim 14 , wherein the plurality of different corpora comprise corpora captured from different time periods, different sources, or a combination thereof.
16 . A system comprising:
storage; processing circuitry configured to execute instructions, that when executed, cause the processing circuitry to: determine a corpus of documents and a model to apply the corpus; iteratively apply the corpus to the model to identify topics in the corpus of documents, each topic comprising a plurality of hierarchically organized components; determine, for each topic, contribution values for the plurality of hierarchically organized components to the topic, wherein each contribution value corresponds to one of the hierarchically organized components, and each contribution value indicates a frequency one of the plurality of hierarchically organized components contributes to the topic; determine a first map comprising the plurality of hierarchically organized components, the contribution values, and the topics, wherein each hierarchically organized component is mapped to a particular topic and has a corresponding contribution value; and store, in the storage, the first map.
17 . The system of claim 16 , comprising the processing circuitry configured to display, the first map in a table format illustrating the plurality of hierarchically organized components and their contribution values to each topic.
18 . The system of claim 16 , comprising the processing circuitry configured to:
generate a second map comprising, for each of the documents of the corpus, a relative distribution of the plurality of hierarchically organized components per topic per document; and store, in the storage, the second map.
19 . The system of claim 18 , comprising determine the relative distribution by summing the contribution value of each hierarchically organized components of the document for each topic.
20 . The system of claim 18 , comprising display, the second map in a table format.
21 . The system of claim 16 , wherein the plurality of hierarchically organized components comprising words.Join the waitlist — get patent alerts
Track US2022318681A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.