Speech-based attention span for voice user interface
Abstract
Techniques for enabling a device to send to a speech processing server further input audio data following a completed utterance dialog to prevent the need for subsequent keywords to be spoken to invoke subsequent commands are described. A system receives input audio data corresponding to an utterance from a device upon the device detecting speech corresponding to a keyword. The system performs speech processing on the input audio data to determine a command. The system determines output data responsive to the command and sends same to the device, thus completing operations regarding the utterance. The system may also send an instruction to the device to: send to the system further input audio data corresponding to further input audio without the device first detecting a wake command.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving first data corresponding to at least one image representing a user; receiving audio data corresponding to an utterance spoken by the user; processing the first data to determine that the utterance is directed at a device; and in response to processing the first data to determine that the utterance is directed at the device, causing speech processing to be performed using the audio data.
22 . The computer-implemented method of claim 21 , further comprising:
processing the first data to determine the user is facing the device.
23 . The computer-implemented method of claim 21 , further comprising:
receiving image data representing the at least one image; and processing the image data using a first component to determine feature data corresponding to the at least one image, wherein the first data includes the feature data, wherein processing the first data to determine that the utterance is directed at a device comprises processing the feature data using at least one classifier.
24 . The computer-implemented method of claim 21 , further comprising:
processing the audio data to determine feature data corresponding to the utterance, wherein processing the first data to determine that the utterance is directed at a device comprises processing the first data and the feature data using at least one classifier.
25 . The computer-implemented method of claim 24 , wherein processing the audio data to determine feature data comprises:
performing automatic speech recognition (ASR) on the audio data to determine ASR result data; and processing the ASR result data to determine the feature data.
26 . The computer-implemented method of claim 21 , wherein causing speech processing to be performed using the audio data comprises sending the audio data to at least one remote device for the speech processing.
27 . The computer-implemented method of claim 21 , wherein audio data was received without detection of a wakeword associated with the utterance.
28 . The computer-implemented method of claim 21 , wherein audio data was received based at least in part on detection of a wakeword associated with the utterance.
29 . The computer-implemented method of claim 21 , further comprising:
processing the first data to determine the user is looking at a second device; and based at least in part on the user looking at the second device, causing output data to be sent to the second device.
30 . The computer-implemented method of claim 29 , wherein determination that the user is looking at the second device occurs after receipt of the audio data.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive first data corresponding to at least one image representing a user;
receive audio data corresponding to an utterance spoken by the user;
process the first data to determine that the utterance is directed at a device; and
in response to processing the first data to determine that the utterance is directed at the device, cause speech processing to be performed using the audio data.
32 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the first data to determine the user is facing the device.
33 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive image data representing the at least one image; and process the image data using a first component to determine feature data corresponding to the at least one image, wherein the first data includes the feature data, wherein the instructions that cause the system to process the first data to determine that the utterance is directed at a device comprise instructions that, when executed by the at least one processor, further cause the system to process the feature data using at least one classifier.
34 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the audio data to determine feature data corresponding to the utterance, wherein the instructions that cause the system to process the first data to determine that the utterance is directed at a device comprise instructions that, when executed by the at least one processor, further cause the system to process the first data and the feature data using at least one classifier.
35 . The system of claim 34 , the instructions that cause the system to process the audio data to determine feature data comprise instructions that, when executed by the at least one processor, further cause the system to:
perform automatic speech recognition (ASR) on the audio data to determine ASR result data; and process the ASR result data to determine the feature data.
36 . The system of claim 31 , wherein the instructions that cause the system to cause speech processing to be performed using the audio data comprise instructions that, when executed by the at least one processor, further cause the system to send the audio data to at least one remote device for the speech processing.
37 . The system of claim 31 , wherein audio data was received without detection of a wakeword associated with the utterance.
38 . The system of claim 31 , wherein audio data was received based at least in part on detection of a wakeword associated with the utterance.
39 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the first data to determine the user is looking at a second device; and based at least in part on the user looking at the second device, cause output data to be sent to the second device.
40 . The system of claim 39 , wherein determination that the user is looking at the second device occurs after receipt of the audio data.Join the waitlist — get patent alerts
Track US2021166686A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.