US2020279552A1PendingUtilityA1
Pre-wakeword speech processing
Est. expiryMar 30, 2035(~8.7 yrs left)· nominal 20-yr term from priority
G10L 25/87G10L 15/08G10L 17/22
62
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system for capturing and processing portions of a spoken utterance command that may occur before a wakeword. The system buffers incoming audio and indicates locations in the audio where the utterance changes, for example when a long pause is detected. When the system detects a wakeword within a particular utterance, the system determines the most recent utterance change location prior to the wakeword and sends the audio from that location to the end of the command utterance to a server for further speech processing.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving first audio data corresponding to first audio; determining, using the first audio data, that a wake expression initiated the first audio; determining the first audio data was directed at a device; causing speech processing to be performed using the first audio data; receiving second audio data corresponding to second audio, wherein the wake expression was absent from the second audio; determining the second audio was directed at the device; and causing speech processing to be performed using the second audio data.
22 . The computer-implemented method of claim 21 , further comprising:
processing the second audio data to determine a beginpoint of the second audio, wherein causing speech processing to be performed using the second audio data comprises causing speech processing to be performed using at least a portion of the second audio data following the beginpoint.
23 . The computer-implemented method of claim 22 , wherein processing the second audio data to determine a beginpoint of the second audio comprises determining a representation of a pause in the second audio data.
24 . The computer-implemented method of claim 21 , further comprising:
processing the first audio data to determine a beginpoint of the first audio, the beginpoint corresponding to the wake expression, wherein causing speech processing to be performed using the first audio data comprises causing speech processing to be performed using at least a portion of the first audio data following the beginpoint.
25 . The computer-implemented method of claim 21 , further comprising storing the second audio data in a temporary storage configured to be overwritten as new audio data is received.
26 . The computer-implemented method of claim 21 , further comprising:
determining first time data corresponding to the second audio data; receiving third audio data corresponding to third audio, wherein the wake expression is represented in the third audio; and determining second time data corresponding to the third audio data, wherein determining the second audio was directed at the device is based at least in part on processing the first time data with regard to the second time data.
27 . The computer-implemented method of claim 21 , further comprising:
detecting an endpoint corresponding to the first audio, wherein the endpoint occurs prior to a beginpoint of the second audio.
28 . The computer-implemented method of claim 21 , further comprising:
receiving third audio data corresponding to third audio, wherein the wake expression is represented in the third audio, wherein determining the second audio was directed at the device is based at least in part on the second audio being part of a same utterance as the third audio.
29 . The computer-implemented method of claim 28 , further comprising:
receiving fourth audio data corresponding to fourth audio, wherein the wake expression was absent from the fourth audio; and causing speech processing to be performed using the fourth audio data based at least in part on the fourth audio being part of the same utterance.
30 . The computer-implemented method of claim 21 , wherein:
the first audio corresponds to a first utterance; and the second audio corresponds to a second utterance different from the first utterance.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive first audio data corresponding to first audio;
determine, using the first audio data, that a wake expression initiated the first audio;
determine the first audio data was directed at a device;
cause speech processing to be performed using the first audio data;
receive second audio data corresponding to second audio, wherein the wake expression was absent from the second audio;
determine the second audio was directed at the device; and
cause speech processing to be performed using the second audio data.
32 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the second audio data to determine a beginpoint of the second audio, wherein the instructions that cause the system to cause speech processing to be performed using the second audio data comprise instructions that cause speech processing to be performed using at least a portion of the second audio data following the beginpoint.
33 . The system of claim 31 , wherein the instructions that cause the system to process the second audio data to determine a beginpoint of the second audio comprise instructions that, when executed by the at least one processor, further cause the system to determine a representation of a pause in the second audio data.
34 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the first audio data to determine a beginpoint of the first audio, the beginpoint corresponding to the wake expression, wherein the instructions that cause the system to cause speech processing to be performed using the first audio data comprise instructions that cause speech processing to be performed using at least a portion of the first audio data following the beginpoint.
35 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
store the second audio data in a temporary storage configured to be overwritten as new audio data is received.
36 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine first time data corresponding to the second audio data; receive third audio data corresponding to third audio, wherein the wake expression is represented in the third audio; and determine second time data corresponding to the third audio data, wherein the instructions that cause the system to determine the second audio was directed at the device are based at least in part on processing the first time data with regard to the second time data.
37 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
detect an endpoint corresponding to the first audio, wherein the endpoint occurs prior to a beginpoint of the second audio.
38 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third audio data corresponding to third audio, wherein the wake expression is represented in the third audio, wherein the instructions that cause the system to determine the second audio was directed at the device are based at least in part on the second audio being part of a same utterance as the third audio.
39 . The system of claim 38 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive fourth audio data corresponding to fourth audio, wherein the wake expression was absent from the fourth audio; and cause speech processing to be performed using the fourth audio data based at least in part on the fourth audio being part of the same utterance.
40 . The system of claim 31 , wherein:
the first audio corresponds to a first utterance; and the second audio corresponds to a second utterance different from the first utterance.Join the waitlist — get patent alerts
Track US2020279552A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.