Visual search for electronic devices
Abstract
The subject technology provides visual search systems and methods that can be used to efficiently perform visual searches on an electronic device. The subject technology provides systems and methods for presenting one or more visual indicators corresponding to the searchable portions of digital content. A visual search may include identifying, at an electronic device, an element of interest in an image. A visual indicator for the element of interest may be overlaid on the image at a location corresponding to the element of interest. The visual indicator may be selectable to cause a display of information associated with the element of interest responsive to the selection.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
obtaining an image provided for display to a user; determining, using a first machine learning model, a plurality of elements present in the image; determining an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image; displaying a visual indicator for the determined element of interest present in the image; based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and in response to detecting a user action with respect to the image, displaying information associated with the element of interest.
2 . The computer-implemented method of claim 1 , wherein determining the element of interest comprises:
determining a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and selecting the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.
3 . The computer-implemented method of claim 2 , wherein the plurality of domains includes one or more of a nature domain, an artwork domain, an animal domain, an automobile domain, a landmark domain, or a media domain.
4 . The computer-implemented method of claim 2 , wherein the first machine learning model comprises an object detection model trained to detect the plurality of elements present in the image and the second machine learning model comprises a classification model trained to determine the respective domain with each respective element of the plurality of elements.
5 . The computer-implemented method of claim 4 , wherein the classification model comprises a gating model.
6 . The computer-implemented method of claim 2 , wherein determining the element of interest from the plurality of elements further comprises determining the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image.
7 . The computer-implemented method of claim 1 , wherein displaying the information associated with the element of interest present in the image comprises displaying a prompt with the visual indicator for the element of interest.
8 . The computer-implemented method of claim 7 , wherein in response to displaying a prompt with the visual indicator for the element of interest, detecting the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest.
9 . The computer-implemented method of claim 8 , wherein in response to detecting the user action:
initiating a search based on the element of interest associated with the visual indicator; and obtaining, from the search, the information associated with the element of interest.
10 . The computer-implemented method of claim 1 , wherein the user action is a tap user action.
11 . The computer-implemented method of claim 1 , further comprising determining, using the first machine learning model, a prediction result indicating whether the element of interest is searchable, wherein the visual indicator for the determined element of interest present in the image is displayed based on the prediction result indicating that the element of interest is searchable.
12 . A device, comprising:
a memory; and a processor configured to:
obtain an image provided for display to a user;
determine, using a first machine learning model, a plurality of elements present in the image;
determine an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image;
display a visual indicator for the determined element of interest present in the image;
based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and
in response to detecting a user action with respect to the image, display information associated with the element of interest.
13 . The device of claim 12 , wherein the processor is configured to:
determine a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and select the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.
14 . The device of claim 12 , wherein the processor is configured to:
determine the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image.
15 . The device of claim 12 , wherein the processor is configured to detect the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest.
16 . The device of claim 15 , wherein the processor is configured to:
initiate a search based on the element of interest associated with the visual indicator; and obtain, from the search, the information associated with the element of interest.
17 . A non-transitory machine-readable medium comprising code that, when executed by a processor, causes the processor to perform a method, comprising:
obtaining an image provided for display to a user; determining, using a first machine learning model, a plurality of elements present in the image; determining an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image; displaying a visual indicator for the determined element of interest present in the image; based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and in response to detecting a user action with respect to the image, displaying information associated with the element of interest in response to detecting a user action with respect to the image.
18 . The non-transitory machine-readable medium of claim 17 , wherein determining the element of interest comprises:
determining a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and selecting the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.
19 . The non-transitory machine-readable medium of claim 17 , wherein determining the element of interest from the plurality of elements further comprises determining the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image.
20 . The non-transitory machine-readable medium of claim 17 , wherein in response to displaying a prompt with the visual indicator for the element of interest, detecting the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest.
21 . The non-transitory machine-readable medium of claim 20 , further comprises:
initiating a search based on the element of interest associated with the visual indicator; and obtaining, from the search, the information associated with the element of interest.Join the waitlist — get patent alerts
Track US12567252B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.