Systems and methods for filtering of computer vision generated tags using natural language processing
Abstract
This disclosure relates to systems, methods, and computer readable media for performing filtering of computer vision generated tags in a media file for the individual user in a multi-format, multi-protocol communication system. One or more media files may be received at a user client. The one or more media files may be automatically analyzed using computer vision models, and computer vision generated tags may be generated in response to analyzing the media file. The tags may then be filtered using Natural Language Processing (NLP) models, and information obtained during NLP tag filtering may be used to train and/or fine-tune one or more of the computer vision models and the NLP models.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A non-transitory computer readable medium comprising computer readable instructions, which, upon execution by one or more processing units, cause the one or more processing units to:
receive a media file for a user, wherein the media file includes one or more objects; automatically analyze the media file using computer vision models responsive to receiving the media file; generate tags for the image responsive to automatically analyzing the media file; filter the tags using Natural Language Processing (NLP) models; and utilize information obtained during filtering of the tags to fine-tune one or more of the computer vision models and the NLP models, wherein the media file includes one of an image or a video.
2 . The non-transitory computer readable medium of claim 1 , wherein the instructions to filter the tags using NLP models further comprise instructions that when executed cause the one or more processing units to select tags that are conceptually closer.
3 . The non-transitory computer readable medium of claim 1 , wherein the instructions to train each of the computer vision models and the NLP models further comprise instructions that when executed cause the one or more processing units to recheck outlier tags in an image corpus for accuracy of the outlier tag.
4 . The non-transitory computer readable medium of claim 1 , wherein the instructions to automatically analyze the media file further comprise instructions that when executed cause the one or more processing units to automatically analyze the media file using one or more of an object segmentation model, object localization model or object detection model.
5 . The non-transitory computer readable medium of claim 1 , wherein the instructions further comprise instructions that when executed cause the one or more processing units to analyze the media file using an object segmentation model for identifying the extent of distinct objects in the image.
6 . The non-transitory computer readable medium of claim 1 , wherein the instructions further comprise instructions that when executed cause the one or more processing units to implement an object detection and recognition model and an object localization model in parallel.
7 . The non-transitory computer readable medium of claim 6 , wherein the instructions further comprise instructions that when executed cause the one or more processing units to implement the object detection and recognition model to determine tags related to general categories of items in the image.
8 . The non-transitory computer readable medium of claim 1 , wherein the instructions further comprise instructions that when executed cause the one or more processing units to implement the object localization model to identify the location of distinct objects in the image.
9 . A system, comprising:
a memory; and one or more processing units, communicatively coupled to the memory, wherein the memory stores instructions to cause the one or more processing units to:
receive an image for a user, wherein the image includes one or more objects;
automatically analyze the image using computer vision models responsive to receiving the media file;
generate tags for the image responsive to automatically analyzing the image;
filter the tags using Natural Language Processing (NLP) models; and
utilize information obtained during filtering of the tags to fine-tune one or more of the computer vision models and the NLP models,
wherein the media file includes one of an image or a video.
10 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to select tags that are conceptually closer responsive to filtering the tags using NLP models.
11 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to recheck outlier tags in an image corpus for accuracy of the outlier tag.
12 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to automatically analyze the image using one or more of an object segmentation model, object localization model or object detection model.
13 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to analyze the media file using an object segmentation model for identifying the extent of distinct objects in the image.
14 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to implement an object detection model and an object localization model in parallel.
15 . The system of claim 14 , the memory further storing instructions to cause the one or more processing units to implement the object detection model to determine tags related to general categories of items in the image.
16 . The system of claim 9 , the memory further storing instructions to cause the one or more processing units to implement the object localization model for identifying the location of distinct objects in the image.
17 . A computer-implemented method, comprising:
receiving an image for a user, wherein the image includes one or more objects; automatically analyzing the image using computer vision models responsive to receiving the media file; generating tags for the image responsive to automatically analyzing the image; filtering the tags using Natural Language Processing (NLP) models; and utilizing information obtained during filtering of the tags to fine-tune one or more of the computer vision models and the NLP models.
18 . The method of claim 17 , further comprising selecting tags that are conceptually closer responsive to filtering the tags.
19 . The method of claim 17 , further comprising rechecking outlier tags in an image corpus for accuracy of the outlier tags.
20 . The method of claim 17 , further comprising automatically analyzing the image using one or more of an object segmentation model, object localization model or object detection model.
21 . The method of claim 17 , further comprising analyzing the media file using an object segmentation model for identifying the extent of distinct objects in the image.
22 . The method of claim 17 , further comprising implementing an object detection model and an object localization model in parallel.
23 . The method of claim 22 , further comprising implementing the object detection model to determine tags related to general categories of items in the image.
24 . The method of claim 17 , further comprising implementing the object localization model to identify a location of distinct objects in the image.
25 . The method of claim 24 , further comprising searching for visually similar objects in a dataset.
26 . The method of claim 21 , further comprising searching for visually similar objects in a dataset.Join the waitlist — get patent alerts
Track US2017193009A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.