US2023253009A1PendingUtilityA1

Hot-word free adaptation of automated assistant function(s)

Assignee: GOOGLE LLCPriority: May 4, 2018Filed: Apr 17, 2023Published: Aug 10, 2023
Est. expiryMay 4, 2038(~11.8 yrs left)· nominal 20-yr term from priority
G10L 15/22G06F 40/20G10L 25/78G06V 40/20G06V 40/18G06V 10/82G06F 40/30G06F 3/167G06V 40/16G06F 3/012G06F 3/013G06F 3/017G06T 7/70G06N 20/00G06T 2207/30196G06V 40/174G06V 10/993G06V 10/235G06F 18/24
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Hot-word free adaptation of one or more function(s) of an automated assistant. Sensor data, from one or more sensor components of an assistant device that provides an automated assistant interface (graphical and/or audible), is processed to determine occurrence and/or confidence metric(s) of various attributes of a user that is proximal to the assistant device. Whether to adapt each of one or more of the function(s) of the automated assistant is based on the occurrence and/or the confidence of one or more of the various attributes. For example, certain processing of at least some of the sensor data can be initiated, such as initiating previously dormant local processing of at least some of the sensor data and/or initiating transmission of at least some of the audio data to remote automated assistant component(s).

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method that facilitates hot-word free interaction between a user and an automated assistant, the method implemented by one or more processors and comprising:
 receiving, at a client device:
 a stream of image frames that are based on output from one or more cameras of the client device, and 
 audio data detected by one or more microphones of the client device; 
   processing, at the client device, the image frames and the audio data to determine co-occurrence of:
 mouth movement of a user, captured by one or more of the image frames, and 
 voice activity of the user; 
   determining, at the client device and based on determining the co-occurrence of the mouth movement of the user and the voice activity of the user, to perform one or both of:
 certain processing of the audio data, and 
 rendering of at least one human perceptible cue via an output component of the client device; and 
   initiating, at the client device, the certain processing of the audio data and/or the rendering of the at least one human perceptible cue, responsive to determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue.   
     
     
         2 . The method of  claim 1 , wherein the certain processing of the audio data is initiated responsive to determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue, and wherein initiating the certain processing of the audio data comprises one or multiple of:
 initiating local automatic speech recognition of the audio data at the client device,   initiating transmission of the audio data to a remote server associated with the automated assistant, and   initiating transmission of recognized text, from the local automatic speech recognition, to the remote server.   
     
     
         3 . The method of  claim 1 , wherein processing, at the client device, the image frames and the audio data to determine co-occurrence of the mouth movement of the user and the voice activity of the user comprises:
 processing both the image frames and the audio data using a locally stored machine learning model trained to distinguish between:
 voice activity that co-occurs with mouth movement and is the result of the mouth movement; and 
 voice activity that is not from the mouth movement, but co-occurs with the mouth movement. 
   
     
     
         4 . The method of  claim 1 , further comprising:
 determining, at the client device, a distance of the user relative to the client device, wherein determining the distance of the user relative to the client device is based on one or both of:
 one or more of the image frames, and 
 additional sensor data from an additional sensor of the client device; 
   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on the distance of the user relative to the client device.   
     
     
         5 . The method of  claim 4 , wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible further based on the distance of the user relative to the client device comprises:
 determining that the distance of the user, relative to the client device satisfies a threshold.   
     
     
         6 . The method of  claim 4 , wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible further based on the distance of the user relative to the client device comprises:
 determining that the distance of the user relative to the client device is closer, to the client device, than one or more previously determined distances of the user relative to the client device.   
     
     
         7 . The method of  claim 1 , further comprising:
 determining, at the client device and based on one or more of the image frames, that a gaze of the user is directed to the client device;   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on determining that the gaze of the user is directed to the client device.   
     
     
         8 . The method of  claim 1 , further comprising:
 determining, at the client device and based on one or more of the image frames, that a body pose of the user is directed to the client device;   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on determining that the body pose of the user is directed to the client device.   
     
     
         9 . A client device comprising:
 at least one camera;   at least one microphone;   at least one display;   one or more processors executing locally stored instructions to:
 receive:
 a stream of image frames that are based on output from the at least one camera, and 
 audio data detected by the at least one microphone; 
 
 process the image frames and the audio data to determine co-occurrence of:
 mouth movement of a user, captured by one or more of the image frames, and 
 voice activity of the user; 
 
 determine, based on determining the co-occurrence of the mouth movement of the user and the voice activity of the user, to perform one or both of:
 certain processing of the audio data, and 
 rendering of at least one human perceptible cue via an output component of the client device; and 
 
 initiate the certain processing of the audio data and/or the rendering of the at least one human perceptible cue, responsive to determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue. 
   
     
     
         10 . The client device of  claim 9 , wherein the certain processing of the audio data is initiated responsive to determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue, and wherein in initiating the certain processing of the audio data one or more of the processors are to:
 initiate local automatic speech recognition of the audio data at the client device,   initiate transmission of the audio data to a remote server associated with the automated assistant, and/or   initiate transmission of recognized text, from the local automatic speech recognition, to the remote server.   
     
     
         11 . The client device of  claim 9 , wherein in processing the image frames and the audio data to determine co-occurrence of the mouth movement of the user and the voice activity of the user, one or more of the processors are to:
 process both the image frames and the audio data using a locally stored machine learning model trained to distinguish between:
 voice activity that co-occurs with mouth movement and is the result of the mouth movement; and 
 voice activity that is not from the mouth movement, but co-occurs with the mouth movement. 
   
     
     
         12 . The client device of  claim 11 , wherein one or more of the processors, in executing the locally stored instructions, are further to:
 determine a distance of the user relative to the client device, wherein determining the distance of the user relative to the client device is based on one or both of:
 one or more of the image frames, and 
 additional sensor data from an additional sensor of the client device; 
   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on the distance of the user relative to the client device.   
     
     
         13 . The client device of  claim 12 , wherein in determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible further based on the distance of the user relative to the client device, one or more of the processors are further to:
 determine that the distance of the user, relative to the client device satisfies a threshold.   
     
     
         14 . The client device of  claim 12 , wherein in determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible further based on the distance of the user relative to the client device, one or more of the processors are to:
 determine that the distance of the user relative to the client device is closer, to the client device, than one or more previously determined distances of the user relative to the client device.   
     
     
         15 . The client device of  claim 11 , wherein one or more of the processors, in executing the locally stored instructions, are further to:
 determine, based on one or more of the image frames, that a gaze of the user is directed to the client device;   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on determining that the gaze of the user is directed to the client device.   
     
     
         16 . The client device of  claim 11 , wherein one or more of the processors, in executing the locally stored instructions, are further to:
 determine, based on one or more of the image frames, that a body pose of the user is directed to the client device;   wherein determining to perform the certain processing of the audio data and/or the rendering of the at least one human perceptible cue is further based on determining that the body pose of the user is directed to the client device.   
     
     
         17 . A method that facilitates hot-word free and touch-free gesture interaction between a user and an automated assistant, the method implemented by one or more processors and comprising:
 receiving, at a client device, a stream of image frames that are based on output from one or more cameras of the client device;   processing, at the client device, the image frames of the stream using at least one trained machine learning model stored locally on the client device to detect occurrence of:
 a gaze of a user that is directed toward the client device; 
   determining, based on detecting the occurrence of the gaze of the user, to generate a response to a gesture of the user that is captured by one or more of the image frames of the stream;   generating the response to the gesture of the user, generating the response comprising:
 determining the gesture of the user based on processing of the one or more of the image frames of the stream, and 
 generating the response based on the gesture of the user and based on content being rendered by the client device at a time of the gesture, wherein generating the response based on the gesture of the user and based on the content being rendered by the client device at the time of the gestures comprises:
 determining that the gesture is assigned to a plurality of responsive actions; 
 selecting, from the plurality of responsive actions assigned to the gesture, a single responsive action, wherein selecting the single responsive action is based on the content being rendered by the client device at the time of the gesture; and 
 generating the response to cause performance of the selected single responsive action, wherein generating the response to cause performance of the selected single responsive action is a result of selecting the single responsive action from the plurality of responsive actions assigned to the gesture; and 
 effectuating the response at the client device. 
 
   
     
     
         18 . The method of  claim 17 , further comprising:
 determining, at the client device, a distance of the user relative to the client device, wherein determining the distance of the user relative to the client device is based on one or both of:
 one or more of the image frames, and 
 additional sensor data from an additional sensor of the client device; and 
   wherein determining to generate the response to the gesture of the user is further based on a magnitude of the distance of the user relative to the client device.   
     
     
         19 . The method of  claim 17 , further comprising:
 determining, based on processing of one or more of the image frames locally at the client device, that the user is a recognized user;   wherein determining to generate the response to the gesture of the user is further based on determining that the user is a recognized user.   
     
     
         20 . The method of  claim 19 , wherein determining to generate the response to the gesture of the user is further based on:
 determining that the user is the recognized user, and
 determining that the same recognized user initiated providing of the content being rendered by the client device.

Join the waitlist — get patent alerts

Track US2023253009A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.