Cascade Architecture for Noise-Robust Keyword Spotting
Abstract
A method (400) includes receiving, at a first processor (110) of a user device (102), streaming multi-channel audio (118) captured by an array of microphones (107), each channel (119) including respective audio features. For each channel, the method also includes processing, by the first processor, using a first stage hotword detector (210), the respective audio features to determine whether a hotword is detected. When the first stage hotword detector detects the hotword, the method also includes the first processor providing chomped raw audio data (212) to a second processor that processes, using a first noise cleaning algorithm (250), the chomped raw audio data to generate a clean monophonic audio chomp (260). The method also includes processing, by the second processor using a second stage hotword detector (220), the clean monophonic audio chomp to detect the hotword.
Claims
exact text as granted — not AI-modified1 . A method comprising:
receiving, at a first processor of a user device, streaming multi-channel audio captured by an array of microphones in communication with the first processor, each channel of the streaming multi-channel audio comprising respective audio features captured by a separate dedicated microphone in the array of microphones; processing, by the first processor, using a first stage hotword detector, the respective audio features of at least one channel of the streaming multi-channel audio to determine whether a hotword is detected by the first stage hotword detector in the streaming multi-channel audio; and when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
providing, by the first processor, chomped multi-channel raw audio data to a second processor of the user device, each channel of the chomped multi-channel raw audio data corresponding to a respective channel of the streaming multi-channel audio and comprising respective raw audio data chomped from the respective channel of the streaming multi-channel audio;
processing, by the second processor, using a first noise cleaning algorithm, each channel of the chomped multi-channel raw audio data to generate a clean monophonic audio chomp;
processing, by the second processor, using a second stage hotword detector, the clean monophonic audio chomp to determine whether the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp; and
when the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp, initiating, by the second processor, a wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.
2 . The method of claim 1 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data comprises an audio segment characterizing the hotword detected by the first stage hotword detector in the streaming multi-channel audio.
3 . The method of claim 2 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data further comprises a prefix segment containing a duration of audio immediately preceding the point in time from when the first stage hotword detector detects the hotword in the streaming multi-channel audio.
4 . The method of claim 1 , wherein:
the second processor operates in a sleep mode when the streaming multi-channel audio is received at the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio is processed by the first processor; and providing the chomped multi-channel audio raw data to the second processor invokes the second processor to transition from the sleep mode to a hotword detection mode.
5 . The method of claim 4 , wherein the second processor executes the first noise cleaning algorithm and the second stage hotword detector while in the hotword detection mode.
6 . The method of claim 1 , further comprising:
processing, by the second processor while processing the clean monophonic audio chomp in parallel, using the second stage hotword detector, the respective raw audio data of one channel of the chomped multi-channel raw audio data to determine whether the hotword is detected by the second stage hotword detector in the respective raw audio data; and when the hotword is detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, initiating, by the second processor, the wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.
7 . The method of claim 6 , further comprising, when the hotword is not detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, preventing, by the second processor initiation of the wake-up process on the user device.
8 . The method of claim 1 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises processing the respective audio features of one channel of the streaming multi-channel audio without canceling noise from the respective audio features.
9 . The method of claim 1 , further comprising:
processing, by the first processor, the respective audio features of each channel of the streaming multi-channel audio to generate a multi-channel cross-correlation matrix; and when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
for each channel of the streaming multi-channel audio, chomping, by the first processor, using the multi-channel cross-correlation matrix, the respective raw audio data from the respective audio features of the respective channel of the streaming multi-channel audio; and
providing, by the first processor, the multi-channel cross-correlation matrix to the second processor,
wherein processing each channel of the chomped multi-channel raw audio data to generate the clean monophonic audio chomp comprises:
computing, using the multi-channel cross-correlation matrix-P provided from the first processor, cleaner filter coefficients for the first noise cleaning algorithm; and
processing, by the first noise cleaning algorithm having the computed cleaner filter coefficients, each channel of the chomped multi-channel raw audio data provided by the first processor to generate the clean monophonic audio chomp.
10 . The method of claim 9 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises:
computing, using the multi-channel cross-correlation matrix, cleaner coefficients for a second noise cleaning algorithm executing on the first processor; processing, by the second noise cleaning algorithm having the computed filter coefficients, each channel of the streaming multi-channel audio to generate a monophonic clean audio stream; and processing, using the first stage hotword detector, the monophonic clean audio stream to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio.
11 . The method of claim 10 , wherein:
the first noise cleaning algorithm applies a first finite impulse response on each channel of the chomped multi-channel raw audio data to generate the chomped monophonic clean audio data, the first FIR comprising a first filter length; and the second noise cleaning algorithm applies a second FIR on each channel of the streaming multi-channel audio to generate the monophonic clean audio stream, the second FIR comprising a second filter length that is less than the first filter length.
12 . The method of claim 1 , wherein:
the first processor comprises a digital signal processor; and the second processor comprises a system on a chip processor.
13 . The method of claim 1 , wherein the user device comprises a rechargeable finite power source, the finite power source powering the first processor and the second processor.
14 . A system comprising:
data processing hardware of a user device, the data processing hardware comprising a first processor and a second processor; and memory processing hardware of the user device, the memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving, at the first processor, streaming multi-channel audio captured by an array of microphones in communication with the first processor, each channel of the streaming multi-channel audio comprising respective audio features captured by a separate dedicated microphone in the array of microphones;
processing, by the first processor, using a first stage hotword detector, the respective audio features of at least one channel of the streaming multi-channel audio to determine whether a hotword is detected by the first stage hotword detector in the streaming multi-channel audio; and
when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
providing, by the first processor, chomped multi-channel raw audio data to the second processor, each channel of the chomped multi-channel raw audio data corresponding to a respective channel of the streaming multi-channel audio and comprising respective raw audio data chomped from the respective channel of the streaming multi-channel audio;
processing, by the second processor, using a first noise cleaning algorithm, each channel of the chomped multi-channel raw audio data to generate a clean monophonic audio chomp;
processing, by the second processor, using a second stage hotword detector, the clean monophonic audio chomp to determine whether the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp; and
when the hotword is detected by the second stage hotword detector in the clean monophonic audio chomp initiating, by the second processor, a wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.
15 . The system of claim 14 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data comprises an audio segment characterizing the hotword detected by the first stage hotword detector in the streaming multi-channel audio.
16 . The system of claim 15 , wherein the respective raw audio data of each channel of the chomped multi-channel raw audio data further comprises a prefix segment containing a duration of audio immediately preceding the point in time from when the first stage hotword detector detects the hotword in the streaming multi-channel audio.
17 . The system of claim 14 , wherein:
the second processor operates in a sleep mode when the streaming multi-channel audio is received at the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio is processed by the first processor; and providing the chomped multi-channel audio raw data to the second processor invokes the second processor to transition from the sleep mode to a hotword detection mode.
18 . The system of claim 17 , wherein the second processor executes the first noise cleaning algorithm and the second stage hotword detector while in the hotword detection mode.
19 . The system of claim 14 , wherein the operations further comprise:
processing, by the second processor while processing the clean monophonic audio chomp in parallel, using the second stage hotword detector, the respective raw audio data of one channel of the chomped multi-channel raw audio data to determine whether the hotword is detected by the second stage hotword detector in the respective raw audio data; and when the hotword is detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, initiating, by the second processor, the wake-up process on the user device for processing the hotword and/or one or more other terms following the hotword in the streaming multi-channel audio.
20 . The system of claim 19 , wherein the operations further comprise, when the hotword is not detected by the second stage hotword detector in either one of the clean monophonic audio chomp or the respective raw audio data, preventing, by the second processor, initiation of the wake-up process on the user device.
21 . The system of claim 14 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises processing the respective audio features of one channel of the streaming multi-channel audio without canceling noise from the respective audio features.
22 . The system of claim 14 , wherein the operations further comprise:
processing, by the first processor, the respective audio features of each channel of the streaming multi-channel audio to generate a multi-channel cross-correlation matrix; and when the first stage hotword detector detects the hotword in the streaming multi-channel audio:
for each channel of the streaming multi-channel audio, chomping, by the first processor using the multi-channel cross-correlation matrix, the respective raw audio data from the respective audio features of the respective channel of the streaming multi-channel audio; and
providing, by the first processor, the multi-channel cross-correlation matrix to the second processor,
wherein processing each channel of the chomped multi-channel raw audio data to generate the clean monophonic audio chomp comprises:
computing, using the multi-channel cross-correlation matrix provided from the first processor cleaner filter coefficients for the first noise cleaning algorithm; and
processing, by the first noise cleaning algorithm having the computed cleaner filter coefficients, each channel of the chomped multi-channel raw audio data provided by the first processor to generate the clean monophonic audio chomp.
23 . The system of claim 22 , wherein processing the respective audio features of the at least one channel of the streaming multi-channel audio to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio comprises:
computing, using the multi-channel cross-correlation matrix, cleaner coefficients for a second noise cleaning algorithm executing on the first processor; processing, by the second noise cleaning algorithm having the computed filter coefficients, each channel of the streaming multi-channel audio to generate a monophonic clean audio stream; and processing, using the first stage hotword detector, the monophonic clean audio stream to determine whether the hotword is detected by the first stage hotword detector in the streaming multi-channel audio.
24 . The system of claim 23 , wherein:
the first noise cleaning algorithm applies a first finite impulse response on each channel of the chomped multi-channel raw audio data to generate the chomped monophonic clean audio data the first FIR comprising a first filter length; and the second noise cleaning algorithm applies a second FIR on each channel of the streaming multi-channel audio to generate the monophonic clean audio stream, the second FIR comprising a second filter length that is less than the first filter length.
25 . The system of claim 14 , wherein:
the first processor comprises a digital signal processor; and the second processor comprises a system on a chip processor.
26 . The system of claim 14 , wherein the user device comprises a rechargeable finite power source, the finite power source powering the first processor and the second processor.Join the waitlist — get patent alerts
Track US2023097197A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.