US2023097197A1PendingUtilityA1

Cascade Architecture for Noise-Robust Keyword Spotting

Assignee: GOOGLE LLCPriority: Apr 8, 2020Filed: Apr 8, 2020Published: Mar 30, 2023
Est. expiryApr 8, 2040(~13.7 yrs left)· nominal 20-yr term from priority
G10L 15/22G06F 3/167G10L 2015/088G10L 21/0208G10L 2021/02166G10L 2015/223G10L 15/10G10L 15/05G10L 21/0216G10L 15/08
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method (400) includes receiving, at a first processor (110) of a user device (102), streaming multi-channel audio (118) captured by an array of microphones (107), each channel (119) including respective audio features. For each channel, the method also includes processing, by the first processor, using a first stage hotword detector (210), the respective audio features to determine whether a hotword is detected. When the first stage hotword detector detects the hotword, the method also includes the first processor providing chomped raw audio data (212) to a second processor that processes, using a first noise cleaning algorithm (250), the chomped raw audio data to generate a clean monophonic audio chomp (260). The method also includes processing, by the second processor using a second stage hotword detector (220), the clean monophonic audio chomp to detect the hotword.

Claims

exact text as granted — not AI-modified
1 . A method comprising:
 receiving, at a first processor of a user device, streaming multi-channel audio captured by an array of microphones in communication with the first processor, each channel of the streaming multi-channel audio comprising respective audio features captured by a separate dedicated microphone in the array of microphones;   processing, by the first processor, using a first stage hotword detector, the respective audio features of at least one channel of the streaming multi-channel audio to determine whether a hotword is detected by the first stage hotword detector in the streaming multi-channel audio; and   when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
 providing, by the first processor, chomped multi-channel raw audio data to a second processor of the user device, each channel of the chomped multi-channel raw audio data corresponding to a respective channel of the streaming multi-channel audio and comprising respective raw audio data chomped from the respective channel of the streaming multi-channel audio; 
 processing, by the second processor, using a first noise cleaning algorithm, each channel of the chomped multi-channel raw audio data to generate a clean monophonic audio chomp; 
 processing, by the second processor, using a second stage hotword detector, the clean monophonic audio chomp to determine whether the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp; and 
 when the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp, initiating, by the second processor, a wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio. 
   
     
     
         2 . The method of  claim 1 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data comprises an audio segment characterizing the hotword detected by the first stage hotword detector in the streaming multi-channel audio. 
     
     
         3 . The method of  claim 2 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data further comprises a prefix segment containing a duration of audio immediately preceding the point in time from when the first stage hotword detector detects the hotword in the streaming multi-channel audio. 
     
     
         4 . The method of  claim 1 , wherein:
 the second processor operates in a sleep mode when the streaming multi-channel audio is received at the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio is processed by the first processor; and   providing the chomped multi-channel audio raw data to the second processor invokes the second processor to transition from the sleep mode to a hotword detection mode.   
     
     
         5 . The method of  claim 4 , wherein the second processor executes the first noise cleaning algorithm and the second stage hotword detector while in the hotword detection mode. 
     
     
         6 . The method of  claim 1 , further comprising:
 processing, by the second processor while processing the clean monophonic audio chomp in parallel, using the second stage hotword detector, the respective raw audio data of one channel of the chomped multi-channel raw audio data to determine whether the hotword is detected by the second stage hotword detector in the respective raw audio data; and   when the hotword is detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, initiating, by the second processor, the wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.   
     
     
         7 . The method of  claim 6 , further comprising, when the hotword is not detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, preventing, by the second processor initiation of the wake-up process on the user device. 
     
     
         8 . The method of  claim 1 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises processing the respective audio features of one channel of the streaming multi-channel audio without canceling noise from the respective audio features. 
     
     
         9 . The method of  claim 1 , further comprising:
 processing, by the first processor, the respective audio features of each channel of the streaming multi-channel audio to generate a multi-channel cross-correlation matrix; and   when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
 for each channel of the streaming multi-channel audio, chomping, by the first processor, using the multi-channel cross-correlation matrix, the respective raw audio data from the respective audio features of the respective channel of the streaming multi-channel audio; and 
 providing, by the first processor, the multi-channel cross-correlation matrix to the second processor, 
   wherein processing each channel of the chomped multi-channel raw audio data to generate the clean monophonic audio chomp comprises:
 computing, using the multi-channel cross-correlation matrix-P provided from the first processor, cleaner filter coefficients for the first noise cleaning algorithm; and 
 processing, by the first noise cleaning algorithm having the computed cleaner filter coefficients, each channel of the chomped multi-channel raw audio data provided by the first processor to generate the clean monophonic audio chomp. 
   
     
     
         10 . The method of  claim 9 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises:
 computing, using the multi-channel cross-correlation matrix, cleaner coefficients for a second noise cleaning algorithm executing on the first processor;   processing, by the second noise cleaning algorithm having the computed filter coefficients, each channel of the streaming multi-channel audio to generate a monophonic clean audio stream; and   processing, using the first stage hotword detector, the monophonic clean audio stream to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio.   
     
     
         11 . The method of  claim 10 , wherein:
 the first noise cleaning algorithm applies a first finite impulse response on each channel of the chomped multi-channel raw audio data to generate the chomped monophonic clean audio data, the first FIR comprising a first filter length; and   the second noise cleaning algorithm applies a second FIR on each channel of the streaming multi-channel audio to generate the monophonic clean audio stream, the second FIR comprising a second filter length that is less than the first filter length.   
     
     
         12 . The method of  claim 1 , wherein:
 the first processor comprises a digital signal processor; and   the second processor comprises a system on a chip processor.   
     
     
         13 . The method of  claim 1 , wherein the user device comprises a rechargeable finite power source, the finite power source powering the first processor and the second processor. 
     
     
         14 . A system comprising:
 data processing hardware of a user device, the data processing hardware comprising a first processor and a second processor; and   memory processing hardware of the user device, the memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
 receiving, at the first processor, streaming multi-channel audio captured by an array of microphones in communication with the first processor, each channel of the streaming multi-channel audio comprising respective audio features captured by a separate dedicated microphone in the array of microphones; 
 processing, by the first processor, using a first stage hotword detector, the respective audio features of at least one channel of the streaming multi-channel audio to determine whether a hotword is detected by the first stage hotword detector in the streaming multi-channel audio; and 
 when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
 providing, by the first processor, chomped multi-channel raw audio data to the second processor, each channel of the chomped multi-channel raw audio data corresponding to a respective channel of the streaming multi-channel audio and comprising respective raw audio data chomped from the respective channel of the streaming multi-channel audio; 
 processing, by the second processor, using a first noise cleaning algorithm, each channel of the chomped multi-channel raw audio data to generate a clean monophonic audio chomp; 
 processing, by the second processor, using a second stage hotword detector, the clean monophonic audio chomp to determine whether the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp; and 
 when the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp initiating, by the second processor, a wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio. 
 
   
     
     
         15 . The system of  claim 14 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data comprises an audio segment characterizing the hotword detected by the first stage hotword detector in the streaming multi-channel audio. 
     
     
         16 . The system of  claim 15 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data further comprises a prefix segment containing a duration of audio immediately preceding the point in time from when the first stage hotword detector detects the hotword in the streaming multi-channel audio. 
     
     
         17 . The system of  claim 14 , wherein:
 the second processor operates in a sleep mode when the streaming multi-channel audio is received at the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio is processed by the first processor; and   providing the chomped multi-channel audio raw data to the second processor invokes the second processor to transition from the sleep mode to a hotword detection mode.   
     
     
         18 . The system of  claim 17 , wherein the second processor executes the first noise cleaning algorithm and the second stage hotword detector while in the hotword detection mode. 
     
     
         19 . The system of  claim 14 , wherein the operations further comprise:
 processing, by the second processor while processing the clean monophonic audio chomp in parallel, using the second stage hotword detector, the respective raw audio data of one channel of the chomped multi-channel raw audio data to determine whether the hotword is detected by the second stage hotword detector in the respective raw audio data; and   when the hotword is detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, initiating, by the second processor, the wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.   
     
     
         20 . The system of  claim 19 , wherein the operations further comprise, when the hotword is not detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, preventing, by the second processor, initiation of the wake-up process on the user device. 
     
     
         21 . The system of  claim 14 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises processing the respective audio features of one channel of the streaming multi-channel audio without canceling noise from the respective audio features. 
     
     
         22 . The system of  claim 14 , wherein the operations further comprise:
 processing, by the first processor, the respective audio features of each channel of the streaming multi-channel audio to generate a multi-channel cross-correlation matrix; and   when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
 for each channel of the streaming multi-channel audio, chomping, by the first processor using the multi-channel cross-correlation matrix, the respective raw audio data from the respective audio features of the respective channel of the streaming multi-channel audio; and 
 providing, by the first processor, the multi-channel cross-correlation matrix to the second processor, 
   wherein processing each channel of the chomped multi-channel raw audio data to generate the clean monophonic audio chomp comprises:
 computing, using the multi-channel cross-correlation matrix provided from the first processor cleaner filter coefficients for the first noise cleaning algorithm; and 
 processing, by the first noise cleaning algorithm having the computed cleaner filter coefficients, each channel of the chomped multi-channel raw audio data provided by the first processor to generate the clean monophonic audio chomp. 
   
     
     
         23 . The system of  claim 22 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises:
 computing, using the multi-channel cross-correlation matrix, cleaner coefficients for a second noise cleaning algorithm executing on the first processor;   processing, by the second noise cleaning algorithm having the computed filter coefficients, each channel of the streaming multi-channel audio to generate a monophonic clean audio stream; and   processing, using the first stage hotword detector, the monophonic clean audio stream to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio.   
     
     
         24 . The system of  claim 23 , wherein:
 the first noise cleaning algorithm applies a first finite impulse response on each channel of the chomped multi-channel raw audio data to generate the chomped monophonic clean audio data the first FIR comprising a first filter length; and   the second noise cleaning algorithm applies a second FIR on each channel of the streaming multi-channel audio to generate the monophonic clean audio stream, the second FIR comprising a second filter length that is less than the first filter length.   
     
     
         25 . The system of  claim 14 , wherein:
 the first processor comprises a digital signal processor; and   the second processor comprises a system on a chip processor.   
     
     
         26 . The system of  claim 14 , wherein the user device comprises a rechargeable finite power source, the finite power source powering the first processor and the second processor.

Join the waitlist — get patent alerts

Track US2023097197A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.