US2024037142A1PendingUtilityA1

Systems and methods for filtering of computer vision generated tags using natural language processing

Assignee: ENTEFY INCPriority: Dec 31, 2015Filed: Aug 10, 2023Published: Feb 1, 2024
Est. expiryDec 31, 2035(~9.4 yrs left)· nominal 20-yr term from priority
G06F 16/5866G06F 16/483G06F 16/51G06F 16/335G06F 16/3344G06F 16/583
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure relates to systems, methods, and computer readable media for performing filtering of computer vision generated tags in a media file for the individual user in a multi-format, multi-protocol communication system. One or more media files may be received at a user client. The one or more media files may be automatically analyzed using computer vision models, and computer vision generated tags may be generated in response to analyzing the media file. The tags may then be filtered using Natural Language Processing (NLP) models, and information obtained during NLP tag filtering may be used to train and/or fine-tune one or more of the computer vision models and the NLP models.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A computer-implemented method comprising:
 determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models;   selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content;   executing a run of the subset of the plurality of computer vision AI models based on the content;   determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value;   filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and   tagging the content in the media file based on the filtered plurality of first computer vision tags.   
     
     
         3 . The computer-implement method of  claim 2 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file. 
     
     
         4 . The computer-implement method of  claim 3 , wherein, prior to the selecting, the computer-implement method further comprises:
 determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.   
     
     
         5 . The computer-implement method of  claim 2 , further comprising:
 identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold;   reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model;   determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and   discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.   
     
     
         6 . The computer-implement method of  claim 2 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags. 
     
     
         7 . The computer-implement method of  claim 6 , wherein, prior to the executing the run, the computer-implemented method further comprises:
 extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and   building at least one scene using the extracted plurality of frames,   wherein the executing the run is further based on the built at least one scene.   
     
     
         8 . The computer-implement method of  claim 2 , wherein the plurality of computer vision AI models comprises at least one of an object segmentation model, an object localization model, an object detection and recognition model, the NLP model, or a relevance feedback loop model. 
     
     
         9 . A system, comprising:
 a non-transitory memory; and   one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations comprising:
 determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models; 
 selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content; 
 executing a run of the subset of the plurality of computer vision AI models based on the content; 
 determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value; 
 filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and 
 tagging the content in the media file based on the filtered plurality of first computer vision tags. 
   
     
     
         10 . The system of  claim 9 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file. 
     
     
         11 . The system of  claim 10 , wherein, prior to the selecting, the operations further comprise:
 determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.   
     
     
         12 . The system of  claim 9 , wherein the operations further comprise:
 identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold;   reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model;   determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and   discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.   
     
     
         13 . The system of  claim 9 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags. 
     
     
         14 . The system of  claim 13 , wherein, prior to the executing the run, the operations further comprise:
 extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and   building at least one scene using the extracted plurality of frames,   wherein the executing the run is further based on the built at least one scene.   
     
     
         15 . The system of  claim 14 , wherein the plurality of computer vision AI models comprises at least one of an object segmentation model, an object localization model, an object detection and recognition model, the NLP model, or a relevance feedback loop model. 
     
     
         16 . A non-transitory computer readable medium comprising computer readable instructions, which, when executed by one or more processing units, cause the one or more processing units to perform operations comprising:
 determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models;   selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content;   executing a run of the subset of the plurality of computer vision AI models based on the content;   determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value;   filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and   tagging the content in the media file based on the filtered plurality of first computer vision tags.   
     
     
         17 . The non-transitory computer readable medium of  claim 16 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file. 
     
     
         18 . The non-transitory computer readable medium of  claim 17 , wherein, prior to the selecting, the operations further comprise:
 determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.   
     
     
         19 . The non-transitory computer readable medium of  claim 16 , wherein the operations further comprise:
 identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold;   reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model;   determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and   discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.   
     
     
         20 . The non-transitory computer readable medium of  claim 16 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags. 
     
     
         21 . The non-transitory computer readable medium of  claim 20 , wherein, prior to the executing the run, the operations further comprise:
 extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and   building at least one scene using the extracted plurality of frames,   wherein the executing the run is further based on the built at least one scene.

Join the waitlist — get patent alerts

Track US2024037142A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.