US2022100965A1PendingUtilityA1
Method and device for adjusting and implementing topic detection processes
Est. expiryJun 7, 2037(~10.9 yrs left)· nominal 20-yr term from priority
G06N 7/01G06F 40/247G06N 20/00G06F 16/355G06F 40/284G06F 16/35G06V 30/416G06F 40/30G06K 9/00469
65
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Aspects of the subject disclosure may include, for example, applying a topic detection process to documents to obtain automatically detected topics and groups of automatically detected words, comparing the automatically detected topics with manually determined topics to determine actual purity metrics, determining an error metric based on a measure of deviation between ideal purity metrics and the actual purity metrics, and adjusting a parameter of the topic detection process according to the error metric resulting in an adjusted topic detection process. Other embodiments are disclosed.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A device comprising:
a processing system including a processor; and a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, comprising: obtaining a first set of topics relating to a document and a first set of word lists, each of the first set of word lists comprising one or more words and corresponding to a respective one of the first set of topics, wherein each topic of the first set of topics is characterized by a probability distribution over a set of words associated with the document, wherein each word of the first set of word lists is in the document, and wherein the obtaining comprises an automated topic detection process using a topic detection parameter; comparing, for each topic of the first set of topics, a corresponding first word list of the first set of word lists with a second word list of a second set of word lists, each of the second set of word lists comprising one or more words and corresponding to a respective one of a second set of topics, the comparing performed according to similarities between respective words of the first word list and respective words of the second word list to determine actual purity metrics, wherein the second set of topics and the second set of word lists are determined in a non-automated process; determining an error metric based on a measure of deviation between ideal purity metrics and the actual purity metrics; and adjusting the topic detection parameter to reduce the error metric resulting in an adjusted topic detection process, wherein the adjusting is performed iteratively, wherein the adjusting is discontinued in accordance with meeting a predetermined criterion.
2 . The device of claim 1 , wherein the ideal purity metrics are based on a determination whether a topic of the first set of topics is a new topic as compared to topics of the second set of topics.
3 . The device of claim 1 , wherein the operations further comprise applying the adjusted topic detection process to the document to obtain an adjusted set of topics and an adjusted set of words that each correspond to one of the adjusted set of topics.
4 . The device of claim 1 , wherein each of the actual purity metrics is determined by:
determining a first similarity quantity by quantifying a largest similarity between first automatically detected words of one of the first set of topics and first determined words of a most similar topic of the second set of topics with respect to the one of the first set of topics; determining a second similarity quantity by quantifying a second largest similarity between the first automatically detected words of the one of the first set of topics and second determined words of a second most similar topic of the second set of topics with respect to the one of the first set of topics; and calculating a differential between the first and second similarity quantities.
5 . The device of claim 1 , wherein the operations further comprise filtering the document prior to performing the topic detection process, wherein the filtering removes particular words, HTML tags, email addresses, or a combination thereof.
6 . The device of claim 1 , wherein the adjusting the parameter of the topic detection process is based on stochastic gradient descent optimization.
7 . The device of claim 1 , wherein the topic detection process comprises a latent dirichlet allocation process.
8 . The device of claim 1 , wherein the measure of deviation comprises a mean squared deviation.
9 . The device of claim 1 , wherein each of the words in the second word list is in the document, wherein the second set of topics and the words in the second word list are derived from a manual analysis of the document.
10 . A method comprising:
obtaining, by a processing system including a processor, a first set of topics relating to a document and a first set of word lists, each of the first set of word lists comprising one or more words and corresponding to a respective one of the first set of topics, wherein each topic of the first set of topics is characterized by a probability distribution over a set of words associated with the document, wherein each word of the first set of word lists is in the document, and wherein the obtaining comprises an automated topic detection process using a topic detection parameter; comparing, by the processing system for each topic of the first set of topics, a corresponding first word list of the first set of word lists with a second word list of a second set of word lists, each of the second set of word lists comprising one or more words and corresponding to a respective one of a second set of topics, the comparing performed according to similarities between respective words of the first word list and respective words of the second word list to determine actual purity metrics, wherein the second set of topics and the second set of word lists are determined in a non-automated process; determining, by the processing system, an error metric based on a measure of deviation between ideal purity metrics and the actual purity metrics; and adjusting, by the processing system, the topic detection parameter to reduce the error metric resulting in an adjusted topic detection process, wherein the adjusting is performed iteratively.
11 . The method of claim 10 , wherein the adjusting is discontinued in accordance with meeting a predetermined criterion.
12 . The method of claim 10 , wherein the ideal purity metrics are based on a determination whether a topic of the first set of topics is a new topic as compared to topics of the second set of topics.
13 . The method of claim 10 , further comprising applying, by the processing system, the adjusted topic detection process to the document to obtain an adjusted set of topics and an adjusted set of words that each correspond to one of the adjusted set of topics.
14 . The method of claim 10 , further comprising filtering, by the processing system, the document prior to performing the topic detection process, wherein the filtering removes particular words, HTML tags, email addresses, or a combination thereof.
15 . The method of claim 10 , wherein the measure of deviation comprises a mean squared deviation.
16 . A non-transitory machine readable medium comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, comprising:
obtaining a first set of topics relating to a document and a first set of word lists, each of the first set of word lists comprising one or more words and corresponding to a respective one of the first set of topics, wherein each topic of the first set of topics is characterized by a probability distribution over a set of words associated with the document, wherein each word of the first set of word lists is in the document, wherein the obtaining comprises an automated topic detection process using a topic detection parameter, and wherein a filtering procedure comprising filtering the document is performed prior to performing the topic detection process; comparing, for each topic of the first set of topics, a corresponding first word list of the first set of word lists with a second word list of a second set of word lists, each of the second set of word lists comprising one or more words and corresponding to a respective one of a second set of topics, the comparing performed according to similarities between respective words of the first word list and respective words of the second word list to determine actual purity metrics, wherein the second set of topics and the second set of word lists are determined in a non-automated process; determining an error metric based on a measure of deviation between ideal purity metrics and the actual purity metrics; and adjusting the topic detection parameter to reduce the error metric resulting in an adjusted topic detection process, wherein the adjusting is performed iteratively, wherein the adjusting is discontinued in accordance with meeting a predetermined criterion.
17 . The non-transitory machine readable medium of claim 16 , wherein the ideal purity metrics are based on a determination whether a topic of the first set of topics is a new topic as compared to topics of the second set of topics.
18 . The non-transitory machine readable medium of claim 16 , wherein the operations further comprise applying the adjusted topic detection process to the document to obtain an adjusted set of topics and an adjusted set of words that each correspond to one of the adjusted set of topics.
19 . The non-transitory machine readable medium of claim 16 , wherein the filtering removes particular words, HTML tags, email addresses, or a combination thereof.
20 . The non-transitory machine readable medium of claim 16 , wherein the measure of deviation comprises a mean squared deviation.Join the waitlist — get patent alerts
Track US2022100965A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.