US2017024664A1PendingUtilityA1

Topic Model Based Media Program Genome Generation For A Video Delivery System

Assignee: HULU LLCPriority: Jan 17, 2014Filed: Sep 30, 2016Published: Jan 26, 2017
Est. expiryJan 17, 2034(~7.5 yrs left)· nominal 20-yr term from priority
G06N 7/01G06F 16/78G06F 40/284G06N 3/126G06F 16/783G06F 16/353G06F 17/30784G06F 17/30707G06N 7/005G06N 99/005G06N 20/00
49
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one embodiment, the method incorporates a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model. Then, the method incorporates a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model. The model is trained with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs. The method further scores terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 incorporating, by a computing device, a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model;   incorporating, by the computing device, a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model;   training, by the computing device, the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and   scoring, by the computing device, terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.   
     
     
         2 . The method of  claim 1 , further comprising incorporating, by the computing device, a parameter that prevents topics from being submerged by other topics that are larger via a third item in the model. 
     
     
         3 . The method of  claim 1 , wherein incorporating, by the computing device, the first set of words that belong to the topic in the set of topics comprises restricting words to a subset of the set of topics via the first item. 
     
     
         4 . The method of  claim 1 , wherein training the model comprises generating the probability distribution when the first set of words are in the topic. 
     
     
         5 . The method of  claim 1 , wherein training the model comprises generating, by the computing device, the probability distribution in view of the relationship between the second set of words that are associated with the topics in the set of topics. 
     
     
         6 . The method of  claim 1 , wherein the relationship specifies two words that co-occur in the topic or do not co-occur in the topic. 
     
     
         7 . The method of  claim 1 , wherein scoring comprises:
 determining, by the computing device, a term frequency for terms found in the textual information for each media program; and   scoring, by the computing device, the topics for each media program based on the probability distribution of terms for the set of topics in the trained model and the corresponding term frequency.   
     
     
         8 . The method of  claim 1 , further comprising normalizing, by the computing device, the scoring of the topics. 
     
     
         9 . The method of  claim 1 , further comprising selecting, by the computing device, a number of highest scored topics for each media program as the genomes for each media program. 
     
     
         10 . The method of  claim 1 , wherein a joint distribution function of a plurality of media programs and the genomes is determined. 
     
     
         11 . The method of  claim 10 , wherein the joint distribution function comprises a joint probability the genome applies to the media program among the plurality of media programs. 
     
     
         12 . The method of  claim 1 , further comprising:
 defining, by the computing device, information for the set of genomes, the set of genomes describing characteristics of media programs; and   defining, by the computing device, which genomes in the set of genomes correspond to which topics in the set of topics.   
     
     
         13 . The method of  claim 1 , wherein training the model comprises:
 inputting, by the computing device, the textual information for the plurality of media programs and the information for the set of genomes into the model.   
     
     
         14 . The method of  claim 1 , wherein training the model comprises:
 determining, by the computing device, words other than the first set of words to associate with the set of topics based on term co-occurrence in the textual information.   
     
     
         15 . A non-transitory computer-readable storage medium containing instructions, that when executed, control a computer system to be configured for:
 incorporating a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model;   incorporating a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model;   training the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and   scoring terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.   
     
     
         16 . The non-transitory computer-readable storage medium of  claim 15 , further configured for incorporating a parameter that prevents topics from being submerged by other topics that are larger via a third item in the model. 
     
     
         17 . The non-transitory computer-readable storage medium of  claim 16 , wherein incorporating the first set of words that belong to the topic in the set of topics comprises restricting words to a subset of the set of topics via the first item. 
     
     
         18 . The non-transitory computer-readable storage medium of  claim 17 , wherein training the model comprises generating the probability distribution when the first set of words are in the topic. 
     
     
         19 . The non-transitory computer-readable storage medium of  claim 15 , wherein training the model comprises generating, by the computing device, the probability distribution in view of the relationship between the second set of words that are associated with the topics in the set of topics. 
     
     
         20 . An apparatus comprising:
 one or more computer processors; and   a non-transitory computer-readable storage medium comprising instructions, that when executed, control the one or more computer processors to be configured for:   incorporating a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model;   incorporating a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model;   training the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and   scoring terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.

Join the waitlist — get patent alerts

Track US2017024664A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.