US2023162728A1PendingUtilityA1

Wakeword detection using a neural network

Assignee: AMAZON TECH INCPriority: Sep 20, 2019Filed: Nov 29, 2022Published: May 25, 2023
Est. expirySep 20, 2039(~13.1 yrs left)· nominal 20-yr term from priority
G10L 15/16G10L 15/063G10L 2015/088G06F 17/15G10L 15/05G06N 3/0442G06N 3/045G06N 3/084
62
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method performs wakeword detection using a feedforward neural network model. A first output of the model indicates when the wakeword appears on a right side of a first window of input audio data. A second output of the model indicates when the wakeword appears in the center of a second window of input audio data. A third output of the model indicates when the wakeword appears on a left side of a third window of input audio data. Using these outputs, the system and method determine a beginpoint and endpoint of the wakeword.

Claims

exact text as granted — not AI-modified
1 .- 20 . (canceled) 
     
     
         21 . A computer-implemented method, comprising:
 receiving audio data representing an utterance;   processing the audio data to determine the audio data represents a first word;   determining first time data representing a beginning of the first word within the audio data;   determining second time data representing an ending of the first word within the audio data;   determining output data comprising the first time data, the second time data, and an indicator of the first word; and   sending the output data.   
     
     
         22 . The computer-implemented method of  claim 21 , wherein:
 the audio data is received by a first device; and   the output data is sent to a second device different from the first device.   
     
     
         23 . The computer-implemented method of  claim 21 , further comprising:
 causing speech processing to be performed using the audio data.   
     
     
         24 . The computer-implemented method of  claim 23 , further comprising:
 sending, from a first device to a second device, second output data corresponding to speech processing results.   
     
     
         25 . The computer-implemented method of  claim 21 , further comprising:
 processing the audio data using a neural network to determine output data; and   determining the first time data using the output data.   
     
     
         26 . The computer-implemented method of  claim 21 , further comprising:
 processing a first portion of the audio data to determine the first portion represents a portion of the first word;   processing a second portion of the audio data to determine the second portion does not represent the first word, wherein the second portion follows the first portion in the audio data;   determining first data representing a boundary between the first portion and the second portion; and   determining the second time data using the first data.   
     
     
         27 . The computer-implemented method of  claim 21 , further comprising:
 processing a first portion of the audio data to determine the first portion does not represent the first word;   processing a second portion of the audio data to determine the second portion represents a portion of the first word, wherein the second portion follows the first portion in the audio data;   determining first data representing a boundary between the first portion and the second portion; and   determining the first time data using the first data.   
     
     
         28 . The computer-implemented method of  claim 21 , further comprising:
 receiving, from at least one microphone, the audio data.   
     
     
         29 . The computer-implemented method of  claim 21 , further comprising:
 determining first data representing a first segment of the audio data;   determining second data representing a second segment of the audio data, wherein the second segment comprises at least a portion of the first segment; and   processing the first data and the second data to determine the audio data represents the first word.   
     
     
         30 . The computer-implemented method of  claim 21 , wherein the first word corresponds to a wakeword. 
     
     
         31 . A system comprising:
 at least one processor; and   at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
 receive audio data representing an utterance; 
 process the audio data to determine the audio data represents a first word; 
 determine first time data representing a beginning of the first word within the audio data; 
 determine second time data representing an ending of the first word within the audio data; 
 determine output data comprising the first time data, the second time data, and an indicator of the first word; and 
 send the output data. 
   
     
     
         32 . The system of  claim 31 , wherein:
 the audio data is received by a first device; and   the output data is sent to a second device different from the first device.   
     
     
         33 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 cause speech processing to be performed using the audio data.   
     
     
         34 . The system of  claim 33 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 send, from a first device to a second device, second output data corresponding to speech processing results.   
     
     
         35 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the audio data using a neural network to determine output data; and   determine the first time data using the output data.   
     
     
         36 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process a first portion of the audio data to determine the first portion represents a portion of the first word;   process a second portion of the audio data to determine the second portion does not represent the first word, wherein the second portion follows the first portion in the audio data;   determine first data representing a boundary between the first portion and the second portion; and   determine the second time data using the first data.   
     
     
         37 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process a first portion of the audio data to determine the first portion does not represent the first word;   process a second portion of the audio data to determine the second portion represents a portion of the first word, wherein the second portion follows the first portion in the audio data;   determine first data representing a boundary between the first portion and the second portion; and   determine the first time data using the first data.   
     
     
         38 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 receive, from at least one microphone, the audio data.   
     
     
         39 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 determine first data representing a first segment of the audio data;   determine second data representing a second segment of the audio data, wherein the second segment comprises at least a portion of the first segment; and   process the first data and the second data to determine the audio data represents the first word.   
     
     
         40 . The system of  claim 31 , wherein the first word corresponds to a wakeword.

Join the waitlist — get patent alerts

Track US2023162728A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.