US2019005332A1PendingUtilityA1
Video understanding platform
Est. expiryDec 30, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06V 10/771G06V 10/761G06V 10/764G06V 20/46H04N 21/44008G06F 18/22G06F 18/2113G06F 18/24G06V 10/44G06K 9/6267H04N 21/4826G06K 9/623G06K 9/4604G06K 9/6215G06K 9/00677G10L 15/26H04L 67/306H04N 21/4394H04N 21/4788G06K 9/00744H04N 21/4532H04N 21/25891G06K 9/00718H04N 21/84H04L 43/045H04L 67/02G06V 20/30G06V 20/41
44
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
In one embodiment, a method includes accessing a video-content object, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object, and determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector. The first type is different from the second type. The method also includes determining a context of the video-content object based on the second feature vector.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
by one or more computing devices, accessing a video-content object; by one or more computing devices, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object; by one or more computing devices, determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and by one or more computing devices, determining a context of the video-content object based on the second feature vector.
2 . The method of claim 1 , wherein:
the first recognition module is an audio-recognition module; the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and the second recognition module is a text-recognition module.
3 . The method of claim 1 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module; the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module; the first recognition module is a text-recognition module and the second recognition module is a video-recognition module; the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module; the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
4 . The method of claim 1 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system; the social graph comprises a plurality of nodes and edges connecting the nodes; and the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.
5 . The method of claim 1 , wherein:
the video-content object comprises frames and audio and is associated with text; and the object in the video-content object is one of:
one or more of the frames;
one or more portions of the audio; or
at least some of the text.
6 . The method of claim 1 , wherein:
the first recognition module is a video-recognition module; the first feature vector represents an intermediate output prediction; and the second recognition module is an audio-recognition module.
7 . The method of claim 1 , further comprising:
by one or more computing devices, determining a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and by one or more computing devices, determining a context of the video-content object based on the third feature vector.
8 . The method of claim 1 , wherein determining the first feature vector comprises:
extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and polling two or more of the first set of feature vectors to generate the first feature vector.
9 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
access a video-content object; determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object; determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and determine a context of the video-content object based on the second feature vector.
10 . The media of claim 9 , wherein:
the first recognition module is an audio-recognition module; the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and the second recognition module is a text-recognition module.
11 . The media of claim 9 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module; the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module; the first recognition module is a text-recognition module and the second recognition module is a video-recognition module; the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module; the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
12 . The media of claim 9 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system; the social graph comprises a plurality of nodes and edges connecting the nodes; and the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.
13 . The media of claim 9 , wherein:
the video-content object comprises frames and audio and is associated with text; and the object in the video-content object is one of:
one or more of the frames;
one or more portions of the audio; or
at least some of the text.
14 . The media of claim 9 , wherein:
the first recognition module is a video-recognition module; the first feature vector represents an intermediate output prediction; and the second recognition module is an audio-recognition module.
15 . The media of claim 9 , wherein the software is further operable when executed to:
determine a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and determine a context of the video-content object based on the third feature vector.
16 . The media of claim 9 , wherein the software is operable to determine the first feature vector by:
extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and polling two or more of the first set of feature vectors to generate the first feature vector.
17 . A system comprising:
one or more processors; and a memory coupled to the processors and comprising instructions operable when executed by the processors to cause the processors to:
access a video-content object;
determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;
determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and
determine a context of the video-content object based on the second feature vector.
18 . The system of claim 17 , wherein:
the first recognition module is an audio-recognition module; the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and the second recognition module is a text-recognition module.
19 . The system of claim 17 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module; the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module; the first recognition module is a text-recognition module and the second recognition module is a video-recognition module; the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module; the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
20 . The method of claim 17 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system; the social graph comprises a plurality of nodes and edges connecting the nodes; and the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.Join the waitlist — get patent alerts
Track US2019005332A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.