US2019220668A1PendingUtilityA1

System and method for sentence directed video object codetection

Assignee: PURDUE RESEARCH FOUNDATIONPriority: Jun 6, 2016Filed: Jun 6, 2017Published: Jul 18, 2019
Est. expiryJun 6, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G06V 10/761G06V 20/47G06V 20/41G06F 18/29G06F 18/22G06T 7/20G06T 2207/10016G06T 2207/10024G06T 7/70G06K 9/6215G06K 9/6296G06K 9/00718G06K 9/726G06V 30/274G06V 20/44
35
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for determining the locations and types of objects in a plurality of videos. The method comprises pairing each video with one or more sentences describing the activity or activities in which those objects participate in the associated video, wherein no use is made of a pretrained object detector. The object locations are specified as rectangles, the object types are specified as nouns, and sentences describe the relative positions and motions of the objects in the videos referred to by the nouns in the sentences. The relative positions and motions of the objects in the video are described by a conjunction of predicates constructed to represent the activity described by the sentences associated with the videos.

Claims

exact text as granted — not AI-modified
1 . A method for determining the locations and types of objects in a plurality of videos, comprising:
 using a computer processor, receiving a plurality of videos;   pairing each of the videos with one or more sentences;   using the processor, describing one or more activities in which those objects participate in a corresponding video; and   wherein no use is made of a pretrained object detector.   
     
     
         2 . The method of  claim 1 , wherein locations of the objects are specified by the processor as rectangles in frames of the videos, the object types are specified as nouns, and sentences describe the relative positions and motions of the objects in the videos referred to by the nouns in the sentences. 
     
     
         3 . The method of  claim 1  or  2 , wherein the relative positions and motions of the objects in the video are described by a conjunction of predicates constructed to represent the activity described by the sentences associated with the videos. 
     
     
         4 . The method according to any previous claim, wherein the locations and types of the objects in the plurality of videos are determined by:
 a. using one or more object proposal mechanisms to propose locations for possible objects in one or more frames of the videos;   b. using one or more object trackers to track the positions of the proposed object locations forward or backward in time;   c. collecting the tracked proposal positions for each proposal into a tube;   d. computing features for each tube based on image features for the portion of the images inside the tubes; or   e. forming a graphical model, wherein:
 i. one or more noun occurrences in sentences associated with a video are associated with vertices in the model; 
 ii. the set of potential labels of each vertex is the set of proposal tubes for the associated video; 
 iii. pairs of vertices that are associated with occurrences of the same noun in two sentences associated with different videos are attached by a binary factor computed as a similarity measure between the tubes selected for from the label sets for the two vertices; 
 iv. collections of vertices that are associated with occurrences of different nouns in the same sentence associated with a video are attached by a factor whose arity is the arity of a predicate in the conjunction of the predicates used to represent the activity described by the sentence where the score of said represents the degree to which the collection of tubes selected for those vertices exhibits the properties of that predicate; or 
 v. the graphical model is solved by selecting a single proposal tube for each vertex from the set of potential labels for that vertex that collectively maximizes a combination of the similarity measure for all pairs of vertices connected by a similarity factor and the predicate scores of all collections of vertices connected by a predicate factor. 
   
     
     
         5 . The method according to any previous claim, wherein the proposal generation mechanism is MCG. 
     
     
         6 . The method according to any previous claim, wherein the proposal generation mechanism is EdgeBoxes. 
     
     
         7 . The method according to any previous claim, wherein the proposals are tracked by CamShift. 
     
     
         8 . The method according to any previous claim, wherein moving proposals are tracked in HSV color space and allowed to change size. 
     
     
         9 . The method according to any previous claim, wherein stationary proposals are tracked in RGB color space and are required to remain of constant size. 
     
     
         10 . The method according to any previous claim, wherein PHOW features are used as image/tube features. 
     
     
         11 . The method according to any previous claim, wherein HOG features are used as image/tube features. 
     
     
         12 . The method according to any previous claim, wherein similarity is measured using a chi-squared distance between image/tube features. 
     
     
         13 . The method according to any previous claim, wherein similarity is measured using Euclidean distance between image/tube features. 
     
     
         14 . The method according to any previous claim, wherein the set of proposal is augmented with proposals rotated by multiples of 90 degrees. 
     
     
         15 . The method according to any previous claim, wherein the similarity measures and predicate scores are combined by summation. 
     
     
         16 . The method according to any previous claim, wherein the similarity measures and predicate scores are combined by taking their product. 
     
     
         17 . The method according to any previous claim, wherein the graphical model is solved using Belief Propagation. 
     
     
         18 . The method according to any previous claim, wherein the set of proposals is augmented by detections produced by a pretrained object detector. 
     
     
         19 . The method of  claim 18 , wherein the method of  claim 1  is first applied and then the method of  claim 18  is applied in one or more subsequent iterations, each iteration using an object detector trained on the proposals selected in earlier iterations.

Join the waitlist — get patent alerts

Track US2019220668A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.