Systems and methods for filtering of computer vision generated tags using natural language processing
Abstract
This disclosure relates to systems, methods, and computer readable media for performing filtering of computer vision generated tags in a media file for the individual user in a multi-format, multi-protocol communication system. One or more media files may be received at a user client. The one or more media files may be automatically analyzed using computer vision models, and computer vision generated tags may be generated in response to analyzing the media file. The tags may then be filtered using Natural Language Processing (NLP) models, and information obtained during NLP tag filtering may be used to train and/or fine-tune one or more of the computer vision models and the NLP models.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A computer-implemented method comprising:
determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models; selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content; executing a run of the subset of the plurality of computer vision AI models based on the content; determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value; filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and tagging the content in the media file based on the filtered plurality of first computer vision tags.
3 . The computer-implement method of claim 2 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file.
4 . The computer-implement method of claim 3 , wherein, prior to the selecting, the computer-implement method further comprises:
determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.
5 . The computer-implement method of claim 2 , further comprising:
identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold; reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model; determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.
6 . The computer-implement method of claim 2 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags.
7 . The computer-implement method of claim 6 , wherein, prior to the executing the run, the computer-implemented method further comprises:
extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and building at least one scene using the extracted plurality of frames, wherein the executing the run is further based on the built at least one scene.
8 . The computer-implement method of claim 2 , wherein the plurality of computer vision AI models comprises at least one of an object segmentation model, an object localization model, an object detection and recognition model, the NLP model, or a relevance feedback loop model.
9 . A system, comprising:
a non-transitory memory; and one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations comprising:
determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models;
selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content;
executing a run of the subset of the plurality of computer vision AI models based on the content;
determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value;
filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and
tagging the content in the media file based on the filtered plurality of first computer vision tags.
10 . The system of claim 9 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file.
11 . The system of claim 10 , wherein, prior to the selecting, the operations further comprise:
determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.
12 . The system of claim 9 , wherein the operations further comprise:
identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold; reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model; determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.
13 . The system of claim 9 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags.
14 . The system of claim 13 , wherein, prior to the executing the run, the operations further comprise:
extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and building at least one scene using the extracted plurality of frames, wherein the executing the run is further based on the built at least one scene.
15 . The system of claim 14 , wherein the plurality of computer vision AI models comprises at least one of an object segmentation model, an object localization model, an object detection and recognition model, the NLP model, or a relevance feedback loop model.
16 . A non-transitory computer readable medium comprising computer readable instructions, which, when executed by one or more processing units, cause the one or more processing units to perform operations comprising:
determining content in a media file using an image analyzer artificial intelligence (AI) model of a plurality of computer vision AI models; selecting a subset of the plurality of computer vision AI models usable to analyze the media file based on the content; executing a run of the subset of the plurality of computer vision AI models based on the content; determining, based on outputs of the subset of the plurality of computer vision AI models from the executed run, a plurality of first computer vision tags for the media file, wherein each of the plurality of first computer vision tags is associated with a confidence value; filtering the plurality of first computer vision tags based on the confidence values and a Natural Language Processing (NLP) model, wherein the filtering removes a portion of the plurality of first computer vision tags based on first corresponding ones of the confidence values at or below a predetermined threshold and prioritizes a remaining portion of the plurality of first computer vision tags based on a ranking of second corresponding ones of the confidence values; and tagging the content in the media file based on the filtered plurality of first computer vision tags.
17 . The non-transitory computer readable medium of claim 16 , wherein the subset of the plurality of computer vision AI models is further selected based on user preferences for a user performing a search associated with the media file.
18 . The non-transitory computer readable medium of claim 17 , wherein, prior to the selecting, the operations further comprise:
determining the user preferences based on at least one of past searches for past content in past media files by the user or ones of the plurality of computer vision AI models usable for identifying the past content for the past searches.
19 . The non-transitory computer readable medium of claim 16 , wherein the operations further comprise:
identifying one of the plurality of first computer vision tags having a corresponding one of the confidence values at or below the predetermined threshold; reprocessing the one of the plurality of first computer vision tags using the subset of the plurality of computer vision AI models and the NLP model; determining that the one of the plurality of first computer vision tags is an irrelevant tag based on the reprocessing; and discarding the one of the plurality of first computer vision tags based on being the irrelevant tag.
20 . The non-transitory computer readable medium of claim 16 , wherein the determining the content includes determining a plurality of second computer vision tags initially used to tag the content in the media file, and wherein the selecting is further based on the plurality of second computer vision tags.
21 . The non-transitory computer readable medium of claim 20 , wherein, prior to the executing the run, the operations further comprise:
extracting a plurality of frames from the media file based on the content and the second plurality of computer vision tags; and building at least one scene using the extracted plurality of frames, wherein the executing the run is further based on the built at least one scene.Join the waitlist — get patent alerts
Track US2024037142A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.