Topic Model Based Media Program Genome Generation For A Video Delivery System
Abstract
In one embodiment, the method incorporates a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model. Then, the method incorporates a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model. The model is trained with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs. The method further scores terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
incorporating, by a computing device, a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model; incorporating, by the computing device, a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model; training, by the computing device, the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and scoring, by the computing device, terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.
2 . The method of claim 1 , further comprising incorporating, by the computing device, a parameter that prevents topics from being submerged by other topics that are larger via a third item in the model.
3 . The method of claim 1 , wherein incorporating, by the computing device, the first set of words that belong to the topic in the set of topics comprises restricting words to a subset of the set of topics via the first item.
4 . The method of claim 1 , wherein training the model comprises generating the probability distribution when the first set of words are in the topic.
5 . The method of claim 1 , wherein training the model comprises generating, by the computing device, the probability distribution in view of the relationship between the second set of words that are associated with the topics in the set of topics.
6 . The method of claim 1 , wherein the relationship specifies two words that co-occur in the topic or do not co-occur in the topic.
7 . The method of claim 1 , wherein scoring comprises:
determining, by the computing device, a term frequency for terms found in the textual information for each media program; and scoring, by the computing device, the topics for each media program based on the probability distribution of terms for the set of topics in the trained model and the corresponding term frequency.
8 . The method of claim 1 , further comprising normalizing, by the computing device, the scoring of the topics.
9 . The method of claim 1 , further comprising selecting, by the computing device, a number of highest scored topics for each media program as the genomes for each media program.
10 . The method of claim 1 , wherein a joint distribution function of a plurality of media programs and the genomes is determined.
11 . The method of claim 10 , wherein the joint distribution function comprises a joint probability the genome applies to the media program among the plurality of media programs.
12 . The method of claim 1 , further comprising:
defining, by the computing device, information for the set of genomes, the set of genomes describing characteristics of media programs; and defining, by the computing device, which genomes in the set of genomes correspond to which topics in the set of topics.
13 . The method of claim 1 , wherein training the model comprises:
inputting, by the computing device, the textual information for the plurality of media programs and the information for the set of genomes into the model.
14 . The method of claim 1 , wherein training the model comprises:
determining, by the computing device, words other than the first set of words to associate with the set of topics based on term co-occurrence in the textual information.
15 . A non-transitory computer-readable storage medium containing instructions, that when executed, control a computer system to be configured for:
incorporating a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model; incorporating a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model; training the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and scoring terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.
16 . The non-transitory computer-readable storage medium of claim 15 , further configured for incorporating a parameter that prevents topics from being submerged by other topics that are larger via a third item in the model.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein incorporating the first set of words that belong to the topic in the set of topics comprises restricting words to a subset of the set of topics via the first item.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein training the model comprises generating the probability distribution when the first set of words are in the topic.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein training the model comprises generating, by the computing device, the probability distribution in view of the relationship between the second set of words that are associated with the topics in the set of topics.
20 . An apparatus comprising:
one or more computer processors; and a non-transitory computer-readable storage medium comprising instructions, that when executed, control the one or more computer processors to be configured for: incorporating a first set of words that belong to a topic in a set of topics that correspond to a set of genomes in a model, the words being incorporated in the model via a first item in the model; incorporating a relationship between a second set of words that are associated with topics in the set of topics via a second item in the model; training the model with respect to the first item and the second item to determine a probability distribution of terms for the set of topics based on analyzing textual information for a plurality of media programs; and scoring terms for each of the plurality of media programs based on the trained model to rank topics that correspond to genomes, the genomes describing characteristics for each media program.Join the waitlist — get patent alerts
Track US2017024664A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.