US12567252B2ActiveUtilityA1

Visual search for electronic devices

Assignee: APPLE INCPriority: Apr 19, 2021Filed: Mar 27, 2024Granted: Mar 3, 2026
Est. expiryApr 19, 2041(~14.7 yrs left)· nominal 20-yr term from priority
H04N 23/617H04N 23/635G06F 18/40G06V 10/25G06Q 30/0625G06N 20/00G06T 11/00G06T 2200/24G06T 11/60G06V 20/20H04N 23/631H04N 23/64
65
PatentIndex Score
0
Cited by
10
References
21
Claims

Abstract

The subject technology provides visual search systems and methods that can be used to efficiently perform visual searches on an electronic device. The subject technology provides systems and methods for presenting one or more visual indicators corresponding to the searchable portions of digital content. A visual search may include identifying, at an electronic device, an element of interest in an image. A visual indicator for the element of interest may be overlaid on the image at a location corresponding to the element of interest. The visual indicator may be selectable to cause a display of information associated with the element of interest responsive to the selection.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method, comprising:
 obtaining an image provided for display to a user;   determining, using a first machine learning model, a plurality of elements present in the image;   determining an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image;   displaying a visual indicator for the determined element of interest present in the image;   based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and   in response to detecting a user action with respect to the image, displaying information associated with the element of interest.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein determining the element of interest comprises:
 determining a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and   selecting the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.   
     
     
         3 . The computer-implemented method of  claim 2 , wherein the plurality of domains includes one or more of a nature domain, an artwork domain, an animal domain, an automobile domain, a landmark domain, or a media domain. 
     
     
         4 . The computer-implemented method of  claim 2 , wherein the first machine learning model comprises an object detection model trained to detect the plurality of elements present in the image and the second machine learning model comprises a classification model trained to determine the respective domain with each respective element of the plurality of elements. 
     
     
         5 . The computer-implemented method of  claim 4 , wherein the classification model comprises a gating model. 
     
     
         6 . The computer-implemented method of  claim 2 , wherein determining the element of interest from the plurality of elements further comprises determining the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image. 
     
     
         7 . The computer-implemented method of  claim 1 , wherein displaying the information associated with the element of interest present in the image comprises displaying a prompt with the visual indicator for the element of interest. 
     
     
         8 . The computer-implemented method of  claim 7 , wherein in response to displaying a prompt with the visual indicator for the element of interest, detecting the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest. 
     
     
         9 . The computer-implemented method of  claim 8 , wherein in response to detecting the user action:
 initiating a search based on the element of interest associated with the visual indicator; and   obtaining, from the search, the information associated with the element of interest.   
     
     
         10 . The computer-implemented method of  claim 1 , wherein the user action is a tap user action. 
     
     
         11 . The computer-implemented method of  claim 1 , further comprising determining, using the first machine learning model, a prediction result indicating whether the element of interest is searchable, wherein the visual indicator for the determined element of interest present in the image is displayed based on the prediction result indicating that the element of interest is searchable. 
     
     
         12 . A device, comprising:
 a memory; and   a processor configured to:
 obtain an image provided for display to a user; 
 determine, using a first machine learning model, a plurality of elements present in the image; 
 determine an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image; 
 display a visual indicator for the determined element of interest present in the image; 
 based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and 
 in response to detecting a user action with respect to the image, display information associated with the element of interest. 
   
     
     
         13 . The device of  claim 12 , wherein the processor is configured to:
 determine a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and   select the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.   
     
     
         14 . The device of  claim 12 , wherein the processor is configured to:
 determine the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image.   
     
     
         15 . The device of  claim 12 , wherein the processor is configured to detect the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest. 
     
     
         16 . The device of  claim 15 , wherein the processor is configured to:
 initiate a search based on the element of interest associated with the visual indicator; and   obtain, from the search, the information associated with the element of interest.   
     
     
         17 . A non-transitory machine-readable medium comprising code that, when executed by a processor, causes the processor to perform a method, comprising:
 obtaining an image provided for display to a user;   determining, using a first machine learning model, a plurality of elements present in the image;   determining an element of interest from the plurality of elements present in the image based in part on a contextual relevance of each element of the plurality of elements present in the image;   displaying a visual indicator for the determined element of interest present in the image;   based on a determination that at least one other element of the plurality of elements present in the image is not searchable, foregoing displaying another visual indicator for the at least one other element; and   in response to detecting a user action with respect to the image, displaying information associated with the element of interest in response to detecting a user action with respect to the image.   
     
     
         18 . The non-transitory machine-readable medium of  claim 17 , wherein determining the element of interest comprises:
 determining a respective domain of a plurality of domains associated with each respective element of the plurality of elements using a second machine learning model that differs from the first machine learning model; and   selecting the element of interest from the plurality of elements present in the image based at least in part on the respective domain determined for each respective element of the plurality of elements.   
     
     
         19 . The non-transitory machine-readable medium of  claim 17 , wherein determining the element of interest from the plurality of elements further comprises determining the element of interest based in part on one or more of a placement, a type, an interaction history, a popularity or a visual relevance of each element in the plurality of elements in the image. 
     
     
         20 . The non-transitory machine-readable medium of  claim 17 , wherein in response to displaying a prompt with the visual indicator for the element of interest, detecting the user action with respect to the visual indicator for the element of interest and responsively displaying the information associated with the element of interest. 
     
     
         21 . The non-transitory machine-readable medium of  claim 20 , further comprises:
 initiating a search based on the element of interest associated with the visual indicator; and   obtaining, from the search, the information associated with the element of interest.

Join the waitlist — get patent alerts

Track US12567252B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.