Systems and methods for adding content to video/multimedia based on metadata
Abstract
An interactive video/multimedia application (IVM application) may specify one or more media assets for playback. The IVM application may define the rendering, composition, and interactivity of one or more the assets, such as video. Video multimedia application data (IVMA data may) be used to define the behavior of the IVM application. The IVMA data may be embodied as a standalone file in a text or binary, compressed format. Alternatively, the IVMA data may be embedded within other media content. A video asset used in the IVM application may include embedded, content-aware metadata that is tightly coupled to the asset. The IVM application may reference the content-aware metadata embedded within the asset to define the rendering and composition of application display elements and user-interactivity features. The interactive video/multimedia application (defined by the video and multimedia application data) may be presented to a viewer in a player application.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An apparatus comprising:
processing logic to automatically detect an object in at least a portion of at least one image frame of video content, and to associate content with the detected object using metadata, wherein the metadata associates at least one multimedia element with the detected object, wherein the metadata comprises a model of the object, and wherein upon rendering of the image frame, the processing logic is to overlay the multimedia element on the detected object in the image frame based on the model; and; and memory coupled to the processing logic, the memory to store the image frame.
2 . The apparatus of claim 1 , wherein to overlay the multimedia element on the detected object, the processing logic uses the model of the object to determine positioning of the multimedia object on the detected object.
3 . The apparatus of claim 1 , wherein the processing logic is further to translate the detected object into the model.
4 . The apparatus of claim 3 , wherein the model is three dimensional.
5 . The apparatus of claim 3 , wherein the model is two dimensional.
6 . The apparatus of claim 1 , wherein the processing logic uses the model to define controls for how the multimedia element is displayed.
7 . The apparatus of claim 1 , wherein the processing logic is further to track movement of the object across multiple image frames and to alter the multimedia element based on the model to display the multimedia element on the object as it moves.
8 . The apparatus of claim 1 , said processing logic to automatically detect an object by detecting at least one of a shape or texture.
9 . A method comprising:
automatically detecting an object in at least a portion of at least one image frame of video content; defining a model of the object; associating content with the detected object using metadata, wherein the metadata associates a location of the object in the frames to the content to provide for metadata-to-content synchronization for the image frame, and wherein the metadata uses the model to define controls for display of the content; and adding the content to the image frame, based on the metadata, at the locations in the image frame, in association with the detected object.
10 . The method of claim 9 , wherein adding content to the image frame comprises using the controls defined from the model of the object to determine positioning of the content on the detected object.
11 . The method of claim 9 , wherein defining the model of the object comprises translating the detected object in the image frame into the model.
12 . The method of claim 11 , wherein the model is three dimensional.
13 . The method of claim 11 , wherein the model is two dimensional.
14 . The method of claim 9 , further comprising tracking movement of the object across multiple image frames and altering the content based on the model to display the multimedia element on the object as it moves.
15 . The method of claim 9 , further comprising detecting an object by detecting at least one of a shape or texture.
16 . One or more non-transitory computer readable media storing instructions to perform a sequence comprising:
automatically detecting an object and its location in a plurality of frames of video while the video is playing; defining, using a model of the object, a set of controls for display of content on the object; associating content with the detected object using metadata, wherein the metadata associates the location of the object in the frames to content to provide for metadata-to-content synchronization for said frames, and upon rendering of the frames, adding the content to the frames, based on the metadata and the set of controls, at the locations in the frames, in association with the detected object.
17 . The media of claim 16 , wherein adding content to the frames comprises using the set of controls defined from the model of the object to determine positioning of the content on the detected object.
18 . The media of claim 16 , further storing instructions to translate the detected object in the frames into the model.
19 . The media of claim 18 , wherein the model is three dimensional.
20 . The media of claim 18 , wherein the model is two dimensional.
21 . The media of claim 16 , further storing instructions to track movement of the object and to display the content on the object as it moves.
22 . The media of claim 16 , further storing instructions to automatically detect an object by detecting at least one of a shape or texture.Join the waitlist — get patent alerts
Track US2019147914A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.