US2021232764A1PendingUtilityA1

Data analytics system and methods for text data

Assignee: AT & T IP I LPPriority: Jul 15, 2016Filed: Apr 16, 2021Published: Jul 29, 2021
Est. expiryJul 15, 2036(~10 yrs left)· nominal 20-yr term from priority
G06F 40/30G06F 16/358G06F 16/285G06F 40/216G06F 16/355
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Aspects of the subject disclosure may include, for example, a process that performs a statistical, natural-language processing analysis on a group of text documents to determine a group of topics. The topics are determined according to parameters obtained by training on a sample of documents. One or more topics in a subset of topics are associated to each document, resulting in topic-document pairs. A bias is identified for each topic-document pair, and clusters of topics are created from the subset of topics. Each cluster of topics is determined from a value for each bias of each topic-document pair and from a frequency of occurrence of each topic. Each cluster is presentable according to a corresponding image configuration based on all or a subset of the bias dimensions and the frequency of occurrence of topics in a cluster that distinguishes the cluster from other clusters. Other embodiments are disclosed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A device, comprising:
 a processing system including a processor; and   a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, the operations comprising:
 performing a first statistical, natural-language processing analysis on a first text document to determine a first plurality of topics based on a first vector of word counts in the first text document; 
 performing a second statistical, natural-language processing analysis on a second text document to determine a second plurality of topics based on a second vector of word counts in the second text document; 
 combining the first plurality of topics and the second plurality of topics to create a plurality of topic-document pairs; 
 identifying a respective bias for each of the plurality of topic-document pairs, resulting in a plurality of biases; and 
 creating a plurality of topic clusters based on the first plurality of topics, the second plurality of topics, and the plurality of biases. 
   
     
     
         2 . The device of  claim 1 , wherein for each of the plurality of topic-document pairs, the respective bias is determined from text in a document identified by that topic-document pair. 
     
     
         3 . The device of  claim 2 , wherein the plurality of topic clusters are determined based on a frequency of occurrence of each topic in respective documents identified by each of the plurality of topic-document pairs. 
     
     
         4 . The device of  claim 3 , wherein the plurality of topic clusters have an image configuration based on the plurality of biases and the frequency of occurrence of each topic that distinguishes one topic cluster of the plurality of topic clusters from another. 
     
     
         5 . The device of  claim 4 , wherein the image configuration comprises size, shape, color coding, or any combination thereof. 
     
     
         6 . The device of  claim 4 , wherein the image configuration comprises an area for each of the plurality of topic clusters. 
     
     
         7 . The device of  claim 6 , wherein a size of the area for each of the plurality of topic clusters represents the frequency of occurrence of each topic in the plurality of topic clusters. 
     
     
         8 . The device of  claim 1 , wherein identifying each of the plurality of biases comprises a latent semantic analysis. 
     
     
         9 . The device of  claim 1 , wherein each of the plurality of biases is identified as one of positive bias, neutral, or negative bias. 
     
     
         10 . The device of  claim 1 , wherein the first statistical, natural-language processing analysis comprises a latent Dirichlet allocation. 
     
     
         11 . The device of  claim 1 , wherein the operations further comprise generating presentable content depicting each of the plurality of topic clusters according to a corresponding image configuration. 
     
     
         12 . The device of  claim 11 , wherein the presentable content includes a Pareto analysis of bias associated with each topic in each of the plurality of topic clusters. 
     
     
         13 . A non-transitory, machine-readable storage medium, comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, the operations comprising:
 performing a first statistical, natural-language processing analysis on a first text document to determine a first plurality of topics based on a first vector of word counts in the first text document;   performing a second statistical, natural-language processing analysis on a second text document to determine a second plurality of topics based on a second vector of word counts in the second text document;   combining the first plurality of topics and the second plurality of topics to create a plurality of topic-document pairs;   identifying a respective bias for each of the plurality of topic-document pairs, resulting in a plurality of biases;   creating a plurality of topic clusters based on the first plurality of topics, the second plurality of topics, and the plurality of biases; and   generating presentable content depicting each of the plurality of topic clusters according to a corresponding image configuration.   
     
     
         14 . The non-transitory, machine-readable storage medium of  claim 13 , wherein for each of the plurality of topic-document pairs, the respective bias is determined from text in a document identified by that topic-document pair. 
     
     
         15 . The non-transitory, machine-readable storage medium of  claim 14 , wherein the plurality of topic clusters are determined based on a frequency of occurrence of each topic in respective documents identified by each of the plurality of topic-document pairs. 
     
     
         16 . The non-transitory, machine-readable storage medium of  claim 14 , wherein the first statistical, natural-language processing analysis comprises a latent Dirichlet allocation. 
     
     
         17 . The non-transitory, machine-readable storage medium of  claim 14 , wherein the presentable content includes a Pareto analysis of bias associated with each topic in each of the plurality of topic clusters. 
     
     
         18 . A method, comprising:
 performing, by a processing system comprising a processor, a first statistical, natural-language processing analysis on a first text document to determine a first plurality of topics based on a first vector of word counts in the first text document;   performing, by the processing system, a second statistical, natural-language processing analysis on a second text document to determine a second plurality of topics based on a second vector of word counts in the second text document;   combining, by the processing system, the first plurality of topics and the second plurality of topics to create a plurality of topic-document pairs;   identifying, by the processing system, a respective bias for each of the plurality of topic-document pairs, resulting in a plurality of biases; and   creating, by the processing system, a plurality of topic clusters based on the first plurality of topics, the second plurality of topics, and the plurality of biases.   
     
     
         19 . The method of  claim 18 , wherein for each of the plurality of topic-document pairs, the respective bias is determined from text in a document identified by that topic-document pair. 
     
     
         20 . The method of  claim 19 , wherein the plurality of topic clusters are determined based on a frequency of occurrence of each topic in respective documents identified by each of the plurality of topic-document pairs.

Join the waitlist — get patent alerts

Track US2021232764A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.