Visual summarization of video for quick understanding
Abstract
The types and locations of particular types of content in a video are visually summarized in a way that facilitates understanding by a viewer. A method may include determining one or more semantic segments of the video. In addition, the method may include determining one or more emotion objects for at least one of the semantic segments. Further, the method may include generating a user interface on a display screen. The user interface may include one window, and in another embodiment, the user interface may include two windows. Moreover, the method may include displaying first indicia of the emotion object in a first window. The horizontal extent of the first window corresponds with the temporal length of the video and the first indicia are displayed at a location corresponding with the temporal appearance of the emotion object in the video.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for rendering a summary of a video, comprising:
determining one or more semantic segments of the video; determining one or more emotion objects for at least one of the semantic segments; generating an interface on a display screen, the interface having a first window; displaying first indicia of the emotion object in the first window, wherein the horizontal extent of the first window corresponds with the temporal length of the video and the first indicia is displayed at a location corresponding with the temporal appearance of the emotion object in the video.
2 . The method of claim 1 , wherein the user interface includes a second window, further comprising displaying a frame of the video in the second window.
3 . The method of claim 1 , further comprising determining a visual object for at least one of the semantic segments and displaying a time line in the first window, the timeline corresponding with the temporal appearance of second object in the video.
4 . The method of claim 1 , further comprising determining an audio object for at least one of the semantic segments and displaying second indicia of the audio object in the first window, wherein the second indicia is displayed at a location corresponding with the temporal rendering of the audio object in the video.
5 . The method of claim 1 , further comprising determining a key word object for at least one of the semantic segments and displaying second indicia of the key word object in the first window, wherein the second indicia is displayed at a location corresponding with the temporal rendering of the key word object in the video.
6 . The method of claim 1 , wherein the first indicia is associated with two or more emotion objects.
7 . A non-transitory computer-readable storage medium having executable code stored thereon to cause a machine to perform a method for rendering a summary of a video, comprising:
determining one or more semantic segments of the video; determining one or more emotion objects for at least one of the semantic segments; generating an interface on a display screen, the interface having a first window; displaying first indicia of the emotion object in the first window, wherein the horizontal extent of the first window corresponds with the temporal length of the video and the first indicia are displayed at a location corresponding with the temporal appearance of the emotion object in the video.
8 . The computer-readable storage medium of claim 7 , wherein the user interface includes a second window, further comprising displaying a frame of the video in the second window.
9 . The computer-readable storage medium of claim 7 , further comprising determining a visual object for at least one of the semantic segments and displaying timeline indicia in the first window, the timeline indicia corresponding with the temporal appearance of second object in the video.
10 . The computer-readable storage medium of claim 7 , further comprising determining an audio object for at least one of the semantic segments and displaying second indicia of the audio object in the first window, wherein the second indicia are displayed at a location corresponding with the temporal rendering of the audio object in the video.
11 . The computer-readable storage medium of claim 7 , further comprising determining a key word object for at least one of the semantic segments and displaying second indicia of the key word object in the first window, wherein the second indicia are displayed at a location corresponding with the temporal rendering of the key word object in the video.
12 . The computer-readable storage medium of claim 7 , wherein the first indicia are associated with two or more emotion objects.Join the waitlist — get patent alerts
Track US2014181668A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.