US2023162728A1PendingUtilityA1
Wakeword detection using a neural network
Est. expirySep 20, 2039(~13.1 yrs left)· nominal 20-yr term from priority
Inventors:Christin JoseYuriy MishchenkoAnish Narendra ShahAlex EscottParind ShahShiv Naga Prasad VitaladevuniThibaud Senechal
G10L 15/16G10L 15/063G10L 2015/088G06F 17/15G10L 15/05G06N 3/0442G06N 3/045G06N 3/084
62
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system and method performs wakeword detection using a feedforward neural network model. A first output of the model indicates when the wakeword appears on a right side of a first window of input audio data. A second output of the model indicates when the wakeword appears in the center of a second window of input audio data. A third output of the model indicates when the wakeword appears on a left side of a third window of input audio data. Using these outputs, the system and method determine a beginpoint and endpoint of the wakeword.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving audio data representing an utterance; processing the audio data to determine the audio data represents a first word; determining first time data representing a beginning of the first word within the audio data; determining second time data representing an ending of the first word within the audio data; determining output data comprising the first time data, the second time data, and an indicator of the first word; and sending the output data.
22 . The computer-implemented method of claim 21 , wherein:
the audio data is received by a first device; and the output data is sent to a second device different from the first device.
23 . The computer-implemented method of claim 21 , further comprising:
causing speech processing to be performed using the audio data.
24 . The computer-implemented method of claim 23 , further comprising:
sending, from a first device to a second device, second output data corresponding to speech processing results.
25 . The computer-implemented method of claim 21 , further comprising:
processing the audio data using a neural network to determine output data; and determining the first time data using the output data.
26 . The computer-implemented method of claim 21 , further comprising:
processing a first portion of the audio data to determine the first portion represents a portion of the first word; processing a second portion of the audio data to determine the second portion does not represent the first word, wherein the second portion follows the first portion in the audio data; determining first data representing a boundary between the first portion and the second portion; and determining the second time data using the first data.
27 . The computer-implemented method of claim 21 , further comprising:
processing a first portion of the audio data to determine the first portion does not represent the first word; processing a second portion of the audio data to determine the second portion represents a portion of the first word, wherein the second portion follows the first portion in the audio data; determining first data representing a boundary between the first portion and the second portion; and determining the first time data using the first data.
28 . The computer-implemented method of claim 21 , further comprising:
receiving, from at least one microphone, the audio data.
29 . The computer-implemented method of claim 21 , further comprising:
determining first data representing a first segment of the audio data; determining second data representing a second segment of the audio data, wherein the second segment comprises at least a portion of the first segment; and processing the first data and the second data to determine the audio data represents the first word.
30 . The computer-implemented method of claim 21 , wherein the first word corresponds to a wakeword.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive audio data representing an utterance;
process the audio data to determine the audio data represents a first word;
determine first time data representing a beginning of the first word within the audio data;
determine second time data representing an ending of the first word within the audio data;
determine output data comprising the first time data, the second time data, and an indicator of the first word; and
send the output data.
32 . The system of claim 31 , wherein:
the audio data is received by a first device; and the output data is sent to a second device different from the first device.
33 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
cause speech processing to be performed using the audio data.
34 . The system of claim 33 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
send, from a first device to a second device, second output data corresponding to speech processing results.
35 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process the audio data using a neural network to determine output data; and determine the first time data using the output data.
36 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process a first portion of the audio data to determine the first portion represents a portion of the first word; process a second portion of the audio data to determine the second portion does not represent the first word, wherein the second portion follows the first portion in the audio data; determine first data representing a boundary between the first portion and the second portion; and determine the second time data using the first data.
37 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
process a first portion of the audio data to determine the first portion does not represent the first word; process a second portion of the audio data to determine the second portion represents a portion of the first word, wherein the second portion follows the first portion in the audio data; determine first data representing a boundary between the first portion and the second portion; and determine the first time data using the first data.
38 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive, from at least one microphone, the audio data.
39 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine first data representing a first segment of the audio data; determine second data representing a second segment of the audio data, wherein the second segment comprises at least a portion of the first segment; and process the first data and the second data to determine the audio data represents the first word.
40 . The system of claim 31 , wherein the first word corresponds to a wakeword.Join the waitlist — get patent alerts
Track US2023162728A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.