US2023053873A1PendingUtilityA1

Invoking automated assistant function(s) based on detected gesture and gaze

Assignee: GOOGLE LLCPriority: May 4, 2018Filed: Nov 4, 2022Published: Feb 23, 2023
Est. expiryMay 4, 2038(~11.8 yrs left)· nominal 20-yr term from priority
G06F 2203/0381G06F 3/013G06F 3/16G10L 15/22G06F 3/017G06F 3/167G06N 20/00G06F 3/0304G06F 3/038
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Invoking one or more previously dormant functions of an automated assistant in response to detecting, based on processing of vision data from one or more vision components: (1) a particular gesture (e.g., of one or more “invocation gestures”) of a user; and/or (2) detecting that a gaze of the user is directed at an assistant device that provides an automated assistant interface (graphical and/or audible) of the automated assistant. For example, the previously dormant function(s) can be invoked in response to detecting the particular gesture, detecting that the gaze of the user is directed at an assistant device for at least a threshold amount of time, and optionally that the particular gesture and the directed gaze of the user co-occur or occur within a threshold temporal proximity of one another.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A client device comprising:
 at least one vision component;   at least one microphone;   one or more processors;   memory operably coupled with the one or more processors, wherein the memory stores instructions that, in response to execution of the instructions by one or more of the processors, cause one or more of the processors to perform the following operations:
 receiving a stream of vision data that is based on output from the vision component of the client device; 
 receiving a stream of audio data that is based on output from the microphone of the client device; 
 determining, based on processing the vision data:
 that a gaze of a user is directed toward the client device, and 
 a user profile for the user; 
 
 determining, based on processing the audio data, that a spoken utterance, included in the audio data:
 temporally corresponds to the gaze, and 
 has voice characteristics that match the user profile that is determined based on processing the vision data; and 
 
 in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the vision data:
 causing at least one dormant function of the automated assistant to be activated. 
 
   
     
     
         2 . The client device of  claim 1 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the vision data comprises:
 transmitting of data, from the client device, to a remote server associated with the automated assistant.   
     
     
         3 . The client device of  claim 1 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the vision data further comprises:
 graphically rendering content that is tailored to the user profile.   
     
     
         4 . The client device of  claim 1 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the vision data comprises:
 automatic speech processing of the audio data.   
     
     
         5 . The client device of  claim 1 , wherein determining, based on processing the vision data, the user profile of the user comprises performing facial recognition based on processing the vision data. 
     
     
         6 . The client device of  claim 1 , wherein determining, based on processing the vision data, that the gaze of the user is directed toward the client device comprises processing the vision data using a trained gaze machine learning model stored locally at the client device. 
     
     
         7 . The client device of  claim 1 , further comprising:
 determining that the user profile is authorized for the client device;   wherein causing the at least one dormant function of the automated assistant to be activated is further contingent on determining that the user profile is authorized for the client device.   
     
     
         8 . A method implemented by one or more processors of a client device that facilitates touch-free interaction between one or more users and an automated assistant, the method comprising:
 processing image frames captured by a camera of the client device;   determining, based on processing the image frames:
 that a gaze of a user is directed toward the client device, and 
 a user profile for the user; 
   processing audio data captured by one or more microphones of the client device;   determining, based on processing the audio data, that a spoken utterance, included in the audio data:
 temporally corresponds to the gaze, and 
 has voice characteristics that match the user profile that is determined based on processing the image frames; and 
   in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the image frames:
 causing at least one dormant function of the automated assistant to be activated. 
   
     
     
         9 . The method of  claim 8 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the image frames comprises:
 transmitting of data, from the client device, to a remote server associated with the automated assistant.   
     
     
         10 . The method of  claim 9 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the image frames comprises:
 automatic speech processing of the audio data.   
     
     
         11 . The method of  claim 8 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the image frames further comprises:
 graphically rendering content that is tailored to the user profile.   
     
     
         12 . The method of  claim 8 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to determining the gaze of the user, and contingent on determining that the spoken utterance temporally corresponds to the gaze and has the voice characteristics that match the user profile that is determined based on processing the image frames comprises:
 automatic speech processing of the audio data.   
     
     
         13 . The method of  claim 8 , wherein determining, based on processing the image frames, the user profile of the user comprises performing facial recognition based on processing at least one of the image frames. 
     
     
         14 . The method of  claim 13 , wherein determining, based on processing the image frames, that the gaze of the user is directed toward the client device comprises processing the image frames using a trained gaze machine learning model stored locally at the client device. 
     
     
         15 . The method of  claim 8 , further comprising:
 determining that the user profile is authorized for the client device;   wherein causing the at least one dormant function of the automated assistant to be activated is further contingent on determining that the user profile is authorized for the client device.   
     
     
         16 . A client device, comprising:
 a vision component;   a presence sensor;   one or more processors, wherein one or more of the processors are configured to:
 detect, based on a signal from the presence sensor, that a human is present in an environment of the presence sensor; 
 in response to detecting that the human is present in the environment:
 activate the vision component to provide a stream of vision data that is based on output from the vision component; 
 
 process the vision data using at least one trained machine learning model stored locally on the client device to monitor for occurrence of both:
 an invocation gesture of a user captured by the vision data, and 
 a gaze of the user that is directed toward the client device; 
 
 detect, based on the monitoring, occurrence of both:
 the invocation gesture, and 
 the gaze; and 
 
   in response to detecting the occurrence of both the invocation gesture and the gaze:
 cause at least one dormant function of an automated assistant to be activated. 
   
     
     
         17 . The client device of  claim 16 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to detecting the occurrence of both the invocation gesture and the gaze comprises:
 transmitting of data, from the client device, to a remote server associated with the automated assistant.   
     
     
         18 . The client device of  claim 16 , wherein the at least one dormant function of the automated assistant, that is caused to be activated in response to detecting the occurrence of both the invocation gesture and the gaze comprises:
 automatic speech processing of the audio data.

Join the waitlist — get patent alerts

Track US2023053873A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.