Query endpointing based on lip detection
Abstract
Systems and methods are described for improving endpoint detection of a voice query submitted by a user. In some implementations, a synchronized video data and audio data is received. A sequence of frames of the video data that includes images corresponding to lip movement on a face is determined. The audio data is endpointed based on first audio data that corresponds to a first frame of the sequence of frames and second audio data that corresponds to a last frame of the sequence of frames. A transcription of the endpointed audio data is generated by an automated speech recognizer. The generated transcription is then provided for output.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A client device, comprising:
a camera; a microphone; a processor; and memory storing instructions that, when executed, cause the processor to:
trigger capturing of video data by the camera;
in response to triggering the capturing of the video data, process the video data to determine which video frames, of the video data, include a face of a user;
process the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech;
in response to determining the video frames are (a) associated with speech:
perform certain processing that is based on audio data that is synchronized with the video data and that is captured via the microphone; and
in response to determining the video frames are (b) associated with an activity other than speech:
bypass performing of the certain processing that is based on the audio data that is synchronized with the video data.
2 . The client device of claim 1 , further comprising a motion sensor, and wherein in triggering capturing of the video data by the camera the processor is to trigger capturing of the video data responsive to detecting motion via the motion sensor.
3 . The client device of claim 1 , wherein in processing the video data to determine which video frames, of the video data, include the face of the user the processor is to use one or more facial recognition techniques.
4 . The client device of claim 1 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data.
5 . The client device of claim 4 , wherein in processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to determine whether the face, in the video frames determined to include the face, include moving lips.
6 . The client device of claim 1 , wherein in processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to process the video frames, determined to include the face of the user, using a deep neural network stored at the client device.
7 . The client device of claim 6 , wherein in processing the video frames, determined to include the face of the user, using the deep neural network, the processor is to:
determine, based on processing the video frames using the deep neural network, a confidence score; and determine, based on whether the confidence score satisfies a threshold, whether the video frames determined to include the face are (a) associated with speech or are (b) associated with the activity other than speech.
8 . The client device of claim 7 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data.
9 . The client device of claim 1 , wherein in determining whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, the processor is to further process the audio data that is synchronized with the video frames.
10 . A method implemented by one or more processors of a client device, the method comprising:
triggering capturing of video data by a camera of the client device; in response to triggering the capturing of the video data:
processing the video data to determine which video frames, of the video data, include a face of a user;
processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech; in response to determining the video frames are (a) associated with speech:
performing certain processing that is based on audio data that is synchronized with the video data and that is captured via a microphone of the client device; and
in response to determining the video frames are (b) associated with an activity other than speech:
bypass performing of the certain processing that is based on the audio data that is synchronized with the video data.
11 . The method of claim 10 , further comprising:
detecting motion via a motion sensor of the client device; wherein triggering capturing of the video data by the camera comprises triggering capturing of the video data responsive to detecting motion via the motion sensor.
12 . The method of claim 10 , wherein processing the video data to determine which video frames, of the video data, include the face of the user comprises using one or more facial recognition techniques.
13 . The method of claim 10 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data.
14 . The method of claim 13 , wherein processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, comprises:
determining whether the face, in the video frames determined to include the face, include moving lips.
15 . The method of claim 10 , wherein processing the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, comprises:
processing the video frames, determined to include the face of the user, using a deep neural network stored at the client device.
16 . The method of claim 15 , wherein processing the video frames, determined to include the face of the user, using the deep neural network, comprises:
determining, based on processing the video frames using the deep neural network, a confidence score; and determining, based on whether the confidence score satisfies a threshold, whether the video frames determined to include the face are (a) associated with speech or are (b) associated with the activity other than speech.
17 . The method of claim 16 , wherein the certain processing that is based on the audio data includes speech recognition on the audio data.
18 . The method of claim 10 , wherein determining whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech, is further based on processing the audio data that is synchronized with the video frames.
19 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor of a client device to:
trigger capturing of video data by a camera of the client device; in response to triggering the capturing of the video data, process the video data to determine which video frames, of the video data, include a face of a user; process the video frames determined to include the face of the user to determine whether the video frames determined to include the face are (a) associated with speech or are (b) associated with an activity other than speech; in response to determining the video frames are (a) associated with speech:
perform certain processing that is based on audio data that is synchronized with the video data and that is captured via a microphone of the client device; and
in response to determining the video frames are (b) associated with an activity other than speech:
bypass performing of the certain processing that is based on the audio data that is synchronized with the video data.Join the waitlist — get patent alerts
Track US2022238112A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.