US2022238112A1PendingUtilityA1

Query endpointing based on lip detection

Assignee: GOOGLE LLCPriority: Mar 14, 2017Filed: Apr 18, 2022Published: Jul 28, 2022
Est. expiryMar 14, 2037(~10.6 yrs left)· nominal 20-yr term from priority
G10L 15/25G10L 15/26G10L 25/78G10L 15/22G10L 21/0356G10L 15/04G10L 15/063G10L 2015/227G06V 40/166G10L 2015/223G10L 15/20G10L 2015/225
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods are described for improving endpoint detection of a voice query submitted by a user. In some implementations, a synchronized video data and audio data is received. A sequence of frames of the video data that includes images corresponding to lip movement on a face is determined. The audio data is endpointed based on first audio data that corresponds to a first frame of the sequence of frames and second audio data that corresponds to a last frame of the sequence of frames. A transcription of the endpointed audio data is generated by an automated speech recognizer. The generated transcription is then provided for output.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A client device, comprising:
 a camera;   a microphone;   a processor; and   memory storing instructions that, when executed, cause the processor to:
 trigger capturing of video data by the camera; 
 in response to triggering the capturing of the video data, process the video data to determine which video frames, of the video data, include a face of a user; 
 process the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech; 
 in response to determining the video frames are (a) associated with speech:
 perform certain processing that is based on audio data that is synchronized with the video data and that is captured via the microphone; and 
 
 in response to determining the video frames are (b) associated with an activity other than speech:
 bypass performing of the certain processing that is based on the audio data that is synchronized with the video data. 
 
   
     
     
         2 . The client device of  claim 1 , further comprising a motion sensor, and wherein in triggering capturing of the video data by the camera the processor is to trigger capturing of the video data responsive to detecting motion via the motion sensor. 
     
     
         3 . The client device of  claim 1 , wherein in processing the video data to determine which video frames, of the video data, include the face of the user the processor is to use one or more facial recognition techniques. 
     
     
         4 . The client device of  claim 1 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data. 
     
     
         5 . The client device of  claim 4 , wherein in processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to determine whether the face, in the video frames determined to include the face, include moving lips. 
     
     
         6 . The client device of  claim 1 , wherein in processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to process the video frames, determined to include the face of the user, using a deep neural network stored at the client device. 
     
     
         7 . The client device of  claim 6 , wherein in processing the video frames, determined to include the face of the user, using the deep neural network, the processor is to:
 determine, based on processing the video frames using the deep neural network, a confidence score; and   determine, based on whether the confidence score satisfies a threshold, whether the video frames determined to include the face are (a) associated with speech or are (b) associated with the activity other than speech.   
     
     
         8 . The client device of  claim 7 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data. 
     
     
         9 . The client device of  claim 1 , wherein in determining whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to further process the audio data that is synchronized with the video frames. 
     
     
         10 . A method implemented by one or more processors of a client device, the method comprising:
 triggering capturing of video data by a camera of the client device;   in response to triggering the capturing of the video data:
 processing the video data to determine which video frames, of the video data, include a face of a user; 
   processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech;   in response to determining the video frames are (a) associated with speech:
 performing certain processing that is based on audio data that is synchronized with the video data and that is captured via a microphone of the client device; and 
   in response to determining the video frames are (b) associated with an activity other than speech:
 bypass performing of the certain processing that is based on the audio data that is synchronized with the video data. 
   
     
     
         11 . The method of  claim 10 , further comprising:
 detecting motion via a motion sensor of the client device;   wherein triggering capturing of the video data by the camera comprises triggering capturing of the video data responsive to detecting motion via the motion sensor.   
     
     
         12 . The method of  claim 10 , wherein processing the video data to determine which video frames, of the video data, include the face of the user comprises using one or more facial recognition techniques. 
     
     
         13 . The method of  claim 10 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data. 
     
     
         14 . The method of  claim 13 , wherein processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, comprises:
 determining whether the face, in the video frames determined to include the face, include moving lips.   
     
     
         15 . The method of  claim 10 , wherein processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, comprises:
 processing the video frames, determined to include the face of the user, using a deep neural network stored at the client device.   
     
     
         16 . The method of  claim 15 , wherein processing the video frames, determined to include the face of the user, using the deep neural network, comprises:
 determining, based on processing the video frames using the deep neural network, a confidence score; and   determining, based on whether the confidence score satisfies a threshold, whether the video frames determined to include the face are (a) associated with speech or are (b) associated with the activity other than speech.   
     
     
         17 . The method of  claim 16 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data. 
     
     
         18 . The method of  claim 10 , wherein determining whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, is further based on processing the audio data that is synchronized with the video frames. 
     
     
         19 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor of a client device to:
 trigger capturing of video data by a camera of the client device;   in response to triggering the capturing of the video data, process the video data to determine which video frames, of the video data, include a face of a user;   process the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech;   in response to determining the video frames are (a) associated with speech:
 perform certain processing that is based on audio data that is synchronized with the video data and that is captured via a microphone of the client device; and 
   in response to determining the video frames are (b) associated with an activity other than speech:
 bypass performing of the certain processing that is based on the audio data that is synchronized with the video data.

Join the waitlist — get patent alerts

Track US2022238112A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.