US2022277351A1PendingUtilityA1

Method and apparatus for labeling data

Assignee: AT & T IP I LPPriority: Dec 17, 2019Filed: May 20, 2022Published: Sep 1, 2022
Est. expiryDec 17, 2039(~13.4 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/0464G06Q 30/0275G06N 3/088G06F 16/355G06N 20/00G06Q 10/107G06N 5/04G06F 40/295G06F 16/93G06Q 30/0269
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Aspects of the subject disclosure may include, for example, determining classes from a corpus based on topic modeling, data clustering and unsupervised learning. Labels are determined for each of the classes and trained models are generated for each of the classes by assignment of a plurality of textual documents to labels based on a highest number of matches. A raw textual document can be tokenized and stop words removed. A corresponding one of the trained models can be selected according to a class that is applicable to subject matter of the raw textual document. The processed document can be assigned to a target label based on a highest number of matches of words. Other embodiments are disclosed.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 receiving, by a processing system including a processor, a raw textual document describing first content that has been added to a media catalog;   processing, by the processing system, the raw textual document to generate a processed document by applying to the raw textual document at least two of: tokenizing, removing stop words, one of bigram and trigram modeling, or name entity recognition analysis;   selecting, by the processing system, a model from a plurality of models according to a class of a plurality of classes, wherein the class is applicable to subject matter of the raw textual document;   automatically assigning, by the processing system, the processed document to a target label via a voting mechanism that counts matches of words in the processed document to labels of the model; and   generating, by the processing system, engagement modeling, cancellation modeling, or both for one or more subscribers according to consumed content that includes the first content and according to the target label.   
     
     
         2 . The method of  claim 1 , further comprising:
 generating, by the processing system, a corpus by processing a plurality of raw textual documents to generate a plurality of textual documents describing content of the media catalog,   wherein the plurality of classes is determined from the corpus.   
     
     
         3 . The method of  claim 2 , further comprising:
 obtaining, by the processing system, the plurality of raw textual documents from Electronic Programming Guide (EPG) data.   
     
     
         4 . The method of  claim 1 , further comprising:
 determining, by the processing system, labels for each class of the plurality of classes using a cosine similarity function.   
     
     
         5 . The method of  claim 4 , wherein the determining of the labels comprises applying a boosting factor during an application of the cosine similarity function. 
     
     
         6 . The method of  claim 5 , wherein the boosting factor is between three to five. 
     
     
         7 . The method of  claim 1 , further comprising:
 generating, by the processing system, a viewer profile for a subscriber according to consumed content that includes the first content and according to the target label.   
     
     
         8 . The method of  claim 1 , further comprising:
 providing the target label to a buyer of electronic advertising in an ad space of the first content responsive to the ad space being presented to a user at an end user device.   
     
     
         9 . The method of  claim 8 , wherein the providing of the target label to the buyer of the electronic advertising is part of a programmatic bidding process. 
     
     
         10 . The method of  claim 1 , wherein the plurality of classes is based on a plurality of textual documents that describes content of the media catalog. 
     
     
         11 . The method of  claim 10 , wherein the content is present in the media catalog when the first content is added to the media catalog. 
     
     
         12 . The method of  claim 10 , wherein multiple documents of the plurality of textual documents are assigned to a single label of a plurality of labels, and wherein the plurality of labels includes the target label. 
     
     
         13 . A device, comprising:
 a processing system including a processor; and   a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, the operations comprising:   generating trained models by assigning each of a plurality of textual documents to a selected label included in a plurality of labels via a voting mechanism that counts matches of words in each document of the plurality of textual documents to the plurality of labels;   receiving and processing a raw textual document to generate a processed document by applying to the raw textual document at least two of: tokenizing, removing stop words, one of bigram and trigram modeling, or name entity recognition analysis, wherein the removing of stop words is based on a type of media content corresponding to first content described by the raw textual document and comprises an identification of characters associated with a type of the raw textual document;   selecting a corresponding model from among the trained models according to a subject matter of the raw textual document;   automatically assigning the processed document to a target label of the corresponding model via the voting mechanism that counts matches of words in the processed document to labels of the corresponding model; and   generating engagement modeling, cancellation modeling, or both for one or more subscribers according to consumed content that includes the first content and according to the target label.   
     
     
         14 . The device of  claim 13 , wherein the plurality of textual documents includes emails, webpage text, articles, transcriptions of recorded voice messages, and closed caption. 
     
     
         15 . The device of  claim 13 , wherein the operations further comprise:
 retrieving the plurality of textual documents from an Electronic Programming Guide (EPG) server.   
     
     
         16 . The device of  claim 13 , wherein the operations further comprise:
 generating a viewer profile for a subscriber according to consumed content that includes media content that the raw textual document describes and according to the target label.   
     
     
         17 . The device of  claim 13 , wherein the operations further comprise:
 providing the target label to a buyer of electronic advertising in an ad space responsive to the ad space being presented to a user at an end user device.   
     
     
         18 . A non-transitory machine-readable medium, comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, the operations comprising:
 processing a raw textual document to generate a processed document by applying Natural Language Processing to the raw textual document, the applying of the Natural Language Processing to the raw textual document comprising at least two of: applying tokenizing, removing stop words, one of bigram and trigram modeling, or name entity recognition analysis, wherein the removing of stop words is based on a type of media content corresponding to first content described by the raw textual document;   automatically assigning the processed document to a target label of a plurality of labels of a model based on a highest number of matches of words in the raw textual document to the target label; and   generating engagement modeling, cancellation modeling, or both for one or more subscribers according to consumed content that includes the first content and according to the target label.   
     
     
         19 . The non-transitory machine-readable medium of  claim 18 , wherein the model is a trained model. 
     
     
         20 . The non-transitory machine-readable medium of  claim 18 , wherein the operations further comprise:
 selecting the model from a plurality of trained models based on a subject matter of the raw textual document.

Join the waitlist — get patent alerts

Track US2022277351A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.