US2019005332A1PendingUtilityA1

Video understanding platform

Assignee: FACEBOOK INCPriority: Dec 30, 2016Filed: Aug 27, 2018Published: Jan 3, 2019
Est. expiryDec 30, 2036(~10.4 yrs left)· nominal 20-yr term from priority
G06V 10/771G06V 10/761G06V 10/764G06V 20/46H04N 21/44008G06F 18/22G06F 18/2113G06F 18/24G06V 10/44G06K 9/6267H04N 21/4826G06K 9/623G06K 9/4604G06K 9/6215G06K 9/00677G10L 15/26H04L 67/306H04N 21/4394H04N 21/4788G06K 9/00744H04N 21/4532H04N 21/25891G06K 9/00718H04N 21/84H04L 43/045H04L 67/02G06V 20/30G06V 20/41
44
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one embodiment, a method includes accessing a video-content object, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object, and determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector. The first type is different from the second type. The method also includes determining a context of the video-content object based on the second feature vector.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 by one or more computing devices, accessing a video-content object;   by one or more computing devices, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;   by one or more computing devices, determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and   by one or more computing devices, determining a context of the video-content object based on the second feature vector.   
     
     
         2 . The method of  claim 1 , wherein:
 the first recognition module is an audio-recognition module;   the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and   the second recognition module is a text-recognition module.   
     
     
         3 . The method of  claim 1 , wherein:
 the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;   the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;   the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or   the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.   
     
     
         4 . The method of  claim 1 , wherein:
 the video-content object corresponds to a node in a social graph of a social-networking system;   the social graph comprises a plurality of nodes and edges connecting the nodes; and   the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.   
     
     
         5 . The method of  claim 1 , wherein:
 the video-content object comprises frames and audio and is associated with text; and   the object in the video-content object is one of:
 one or more of the frames; 
 one or more portions of the audio; or 
 at least some of the text. 
   
     
     
         6 . The method of  claim 1 , wherein:
 the first recognition module is a video-recognition module;   the first feature vector represents an intermediate output prediction; and   the second recognition module is an audio-recognition module.   
     
     
         7 . The method of  claim 1 , further comprising:
 by one or more computing devices, determining a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and   by one or more computing devices, determining a context of the video-content object based on the third feature vector.   
     
     
         8 . The method of  claim 1 , wherein determining the first feature vector comprises:
 extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and   polling two or more of the first set of feature vectors to generate the first feature vector.   
     
     
         9 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
 access a video-content object;   determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;   determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and   determine a context of the video-content object based on the second feature vector.   
     
     
         10 . The media of  claim 9 , wherein:
 the first recognition module is an audio-recognition module;   the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and   the second recognition module is a text-recognition module.   
     
     
         11 . The media of  claim 9 , wherein:
 the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;   the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;   the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or   the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.   
     
     
         12 . The media of  claim 9 , wherein:
 the video-content object corresponds to a node in a social graph of a social-networking system;   the social graph comprises a plurality of nodes and edges connecting the nodes; and   the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.   
     
     
         13 . The media of  claim 9 , wherein:
 the video-content object comprises frames and audio and is associated with text; and   the object in the video-content object is one of:
 one or more of the frames; 
 one or more portions of the audio; or 
 at least some of the text. 
   
     
     
         14 . The media of  claim 9 , wherein:
 the first recognition module is a video-recognition module;   the first feature vector represents an intermediate output prediction; and   the second recognition module is an audio-recognition module.   
     
     
         15 . The media of  claim 9 , wherein the software is further operable when executed to:
 determine a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and   determine a context of the video-content object based on the third feature vector.   
     
     
         16 . The media of  claim 9 , wherein the software is operable to determine the first feature vector by:
 extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and   polling two or more of the first set of feature vectors to generate the first feature vector.   
     
     
         17 . A system comprising:
 one or more processors; and   a memory coupled to the processors and comprising instructions operable when executed by the processors to cause the processors to:
 access a video-content object; 
 determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object; 
 determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and 
 determine a context of the video-content object based on the second feature vector. 
   
     
     
         18 . The system of  claim 17 , wherein:
 the first recognition module is an audio-recognition module;   the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and   the second recognition module is a text-recognition module.   
     
     
         19 . The system of  claim 17 , wherein:
 the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;   the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;   the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;   the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or   the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.   
     
     
         20 . The method of  claim 17 , wherein:
 the video-content object corresponds to a node in a social graph of a social-networking system;   the social graph comprises a plurality of nodes and edges connecting the nodes; and   the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.

Join the waitlist — get patent alerts

Track US2019005332A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.