Context-aware event based annotation system for media asset
Abstract
Various embodiments described herein support or provide for annotation of a media asset, such as an audio asset or a video asset, based on one or more events identified within content of the media asset. In particular, some embodiments can determine one or more of the following details with respect to content of a given media asset, which can represent annotations that enable determination of contextual information for the given media asset: events; event classification labels for events; subclassifications labels for events; scenes comprising events; attributes of scenes; themes presented by the content; and title-level attributes of the given media asset.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method comprising:
accessing, by a hardware processor, content data for a current media asset;
determining, by the hardware processor, a set of events within the content data by scanning the content data for events that relate to at least one event classification, each event in the set of events comprising at least one of a visual content element, a textual content feature, or an audio content feature from the content data being presented at a timestamp of the current media asset, each event in the set of events being associated with an event classification label selected from a predetermined event classification library, the predetermined event classification library comprising a plurality of event classification labels where each event classification label comprises a set of available event subclassification labels;
for each individual event in the set of events and based on an identified event classification label of the individual event, determining a set of identified event subclassification labels for the individual event, the set of identified event subclassification labels being selected from the set of available subclassification labels for the identified event classification as provided by the predetermined event classification library, at least one identified event subclassification label in the set of identified event subclassification labels describing at least:
an intent of a context of the individual event; and
an outcome of the context of the individual event;
determining, by the hardware processor, a set of scenes within the content data, each scene in the set of scenes comprising a subset of events from the set of events;
for each individual scene in the set of scenes, determining a set of scene attributes for the individual scene;
determining, by the hardware processor, a set of themes for the current media asset based on at least one of the set of scenes or the set of scene attributes;
determining, by the hardware processor, a set of title attributes for the current media asset based on at least the set of themes and metadata associated with the media asset; and
generating, by the hardware processor, contextual data for the current media asset based on at least one of the set of events, the set of event classification labels determined for the set of events, the sets of identified event subclassification labels determined for the set of events, the set of scenes, sets of scene attributes for the set of scenes, the set of themes, or the set of title attributes.
2. The method of claim 1 , wherein the predetermined event classification library is configured such that event classification labels and event subclassification labels of the predetermined event classification library cause events in the current media asset to be classified without cultural bias.
3. The method of claim 1 , wherein an individual event subclassification label determined for the individual event provides detail with respect to a context of the individual event.
4. The method of claim 3 , wherein the individual event subclassification label describes at least: a description of the context of the individual event; an explanation of the context of the individual event; an how the context of the individual event is presented in the content data of the current media asset.
5. The method of claim 1 , comprising:
causing, by the hardware processor, display of a graphical user interface for screening the current media asset, the graphical user interface being configured to receive user input that identifies at least one of:
one or more events in the content data of the current media asset;
one or more event classification labels for an event of the current media asset;
one or more event subclassification labels for an event of the current media;
one or more scenes in the content data of the current media asset;
one or more themes for the current media asset; or
one or more title attributes for the current media asset.
6. The method of claim 1 , wherein the scanning of the media asset is performed using an event scanner; and wherein the event scanner comprises a machine learning model trained to automatically identify a select event at a select timestamp of the current media asset based on a set of signals provided by at least one computer vision analysis, audio analysis, or natural language processing of content presented by the current media asset at the select timestamp.
7. The method of claim 6 , wherein the machine learning model is trained based on contextual data of another media asset.
8. The method of claim 6 , comprising:
causing the machine learning model to further train using at least a portion of the generated contextual data.
9. The method of claim 1 , wherein the determining the set of identified event subclassification labels for the individual event is performed using an event classifier, wherein the individual event is at a select timestamp of the current media asset; and wherein the event classifier comprises a machine learning model trained to automatically identify the set of identified event subclassification labels for the individual event based on a set of signals provided by at least one computer vision analysis, audio analysis, or natural language processing of content presented by the current media asset at the select timestamp.
10. The method of claim 9 , wherein the machine learning model is trained based on contextual data of another media asset.
11. The method of claim 9 , comprising:
causing the machine learning model to further train using at least a portion of the generated contextual data.
12. The method of claim 1 , wherein the determining the set of scene attributes for the individual scene is performed using a scene analyzer; and wherein the scene analyzer comprises a machine learning model trained to automatically identify the set of scene attributes for the individual scene based on at least one of:
one or more events of the individual scene;
one or more event classification labels for the one or more events; or
one or more event subclassification labels for the one or more events.
13. The method of claim 12 , wherein the machine learning model is trained based on contextual data of another media asset.
14. The method of claim 12 , comprising:
causing the machine learning model to further train using at least portion of the generated contextual data.
15. The method of claim 1 , wherein the determining the set of scene attributes for the individual scene comprises determining at least one of:
determining a frequency of events in the individual scene;
determining a mixture of events, with different event classification labels, in the individual scene;
determining a time distance between events in the individual scene; or
determining a duration of the individual scene.
16. The method of claim 1 , wherein the determining of the set of themes for the current media asset is performed using a theme analyzer; and wherein the theme analyzer comprises a machine learning model trained to automatically identify the set of themes for the current media asset based on at least one of the set of scenes or the set of scene attributes.
17. The method of claim 16 , wherein the machine learning model is trained based on contextual data of another media asset.
18. The method of claim 16 , comprising:
causing the machine learning model to further train using at least a portion of the generated contextual data.
19. The method of claim 1 , comprising:
causing, by the hardware processor, display of a graphical user interface for screening the current media asset, the graphical user interface including a time bar for the content data of the current media asset, and the time bar including a visual indicator for each timestamp of the current media asset that is associated with a select event from the set of events or a select scene from the set of scenes.
20. The method of claim 1 , comprising:
causing, by the hardware processor, display of a graphical user interface for screening the current media asset, the graphical user interface including a listing of tags that correspond to events from the set of events or scenes from the set of scenes.
21. The method of claim 20 , wherein in the listing of tags, a select tag for a select event or a select scene is displayed with one or more event classification labels of the select event or the select scene.
22. The method of claim 1 , comprising:
causing, by the hardware processor, a media software tool configured to process the current media asset based on the contextual data for the current media asset.
23. The method of claim 1 , wherein the metadata comprises at least one of: an attribute describing a genre of the current media asset; an attribute describing how the content data of the current media asset is presented; an attribute describing a cast or a crew member listed for the current media asset; an attribute describing entities involved in production of the current media asset; an attribute describing a production or release date for the current media asset; or a runtime of the current media asset.
24. A system comprising:
a memory storing instructions; and
one or more hardware processors communicatively coupled to the memory and configured by the instructions to perform operations comprising:
accessing content data for a current media asset;
determining a set of events within the content data by scanning the content data for events that relate to at least one event classification, each event in the set of events comprising at least one of a visual content element, a textual content feature, or an audio content feature from the content data being presented at a timestamp of the current media asset, each event in the set of events being associated with an event classification label selected from a predetermined event classification library, the predetermined event classification library comprising a plurality of event classification labels where each event classification label comprises a set of available event subclassification labels;
for each individual event in the set of events and based on an identified event classification label of the individual event, determining a set of identified event subclassification labels for the individual event, the set of identified event subclassification labels being selected from the set of available subclassification labels for the identified event classification as provided by the predetermined event classification library, at least one identified event subclassification label in the set of identified event subclassification labels describing at least:
an intent of a context of the individual event; and
an outcome of the context of the individual event;
determining a set of scenes within the content data, each scene in the set of scenes comprising a subset of events from the set of events;
for each individual scene in the set of scenes, determining a set of scene attributes for the individual scene;
determining a set of themes for the current media asset based on at least one of the set of scenes or the set of scene attributes;
determining a set of title attributes for the current media asset based on at least the set of themes and metadata associated with the media asset; and
generating contextual data for the current media asset based on at least one of the set of events, the set of event classification labels determined for the set of events, the sets of identified event subclassification labels determined for the set of events, the set of scenes, sets of scene attributes for the set of scenes, the set of themes, or the set of title attributes.
25. A non-transitory computer-readable medium comprising instructions that, when executed by a hardware processor of a device, cause the device to perform operations comprising:
accessing content data for a current media asset;
determining a set of events within the content data by scanning the content data for events that relate to at least one event classification, each event in the set of events comprising at least one of a visual content element, a textual content feature, or an audio content feature from the content data being presented at a timestamp of the current media asset, each event in the set of events being associated with an event classification label selected from a predetermined event classification library, the predetermined event classification library comprising a plurality of event classification labels where each event classification label comprises a set of available event subclassification labels;
for each individual event in the set of events and based on an identified event classification label of the individual event, determining a set of identified event subclassification labels for the individual event, the set of identified event subclassification labels being selected from the set of available subclassification labels for the identified event classification as provided by the predetermined event classification library, at least one identified event subclassification label in the set of identified event subclassification labels describing at least:
an intent of a context of the individual event; and
an outcome of the context of the individual event;
determining a set of scenes within the content data, each scene in the set of scenes comprising a subset of events from the set of events;
for each individual scene in the set of scenes, determining a set of scene attributes for the individual scene;
determining a set of themes for the current media asset based on at least one of the set of scenes or the set of scene attributes;
determining a set of title attributes for the current media asset based on at least the set of themes and metadata associated with the media asset; and
generating contextual data for the current media asset based on at least one of the set of events, the set of event classification labels determined for the set of events, the sets of identified event subclassification labels determined for the set of events, the set of scenes, sets of scene attributes for the set of scenes, the set of themes, or the set of title attributes.Join the waitlist — get patent alerts
Track US11776261B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.