Multi-modal audio processing for voice-controlled devices
Abstract
A voice-controlled device includes a microphone to receive a set of sound waves that includes speech uttered by a user and other sound, and to output a first audio signal that includes a contribution from the speech uttered by the user and a contribution from the other sound. The device also includes a receiver to receive an electromagnetic signal and to output a second audio signal obtained from the electromagnetic signal. An audio pre-processor of the device processes the first audio signal using the second audio signal to reduce the contribution from the other sound in a processed audio signal. The voice-controlled device then provides the processed audio signal to a speech recognition module to determine a voice command issued by the user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of processing an audio signal for a voice-controlled device, the method comprising:
receiving a set of sound waves at a microphone of the voice-controlled device, the set of sound waves comprising speech uttered by a user and other sound; converting, using the microphone, the set of sound waves into a first audio signal that includes a contribution from the speech uttered by the user and a contribution from the other sound; receiving, at a receiver of the voice-controlled device, an electromagnetic signal; obtaining a second audio signal from the electromagnetic signal; processing the first audio signal using the second audio signal to reduce the contribution from the other sound in a processed audio signal; and performing speech recognition on the processed audio signal to determine a voice command issued by the user.
2 . The method of claim 1 , further comprising:
determining an amplitude of a version of the second audio signal that is present within the first audio signal; scaling the second audio signal obtained from the electromagnetic signal based on the determined amplitude to generate a modified version of the second audio signal; and subtracting the modified version of the second audio signal from the first audio signal as at least a part of said processing.
3 . The method of claim 1 , further comprising:
determining a time difference between a version of the second audio signal that is present within the first audio signal and the second audio signal that is obtained from the electromagnetic signal; delaying the second audio signal obtained from the electromagnetic signal using the determined time difference to generate a modified version of the second audio signal; and subtracting the modified version of the second audio signal from the first audio signal as at least a part of said processing.
4 . The method of claim 1 , further comprising:
evaluating a cross-correlation function between the first audio signal and the second audio signal; obtaining a time delay and/or a scaling factor from an output of the cross-correlation function; and applying the time delay and/or the scaling factor to the second audio signal to obtain a modified version of the second audio signal; and subtracting the modified version of the second audio signal from the first audio signal as at least a part of said processing.
5 . The method of claim 1 , wherein the electromagnetic signal comprises one or more electromagnetic signals, the method further comprising:
obtaining a plurality of other audio signals from the one or more electromagnetic signals, the plurality of other audio signals including the second audio signal; detecting one or more of the plurality of other audio signals within the first audio signal; and subtracting versions of the detected one or more of the plurality of other audio signals from the first audio signal as at least a part of said processing.
6 . The method of claim 5 , wherein the one or more electromagnetic signals comprise at least one modulated radio signal and the plurality of other audio signals are obtained by demodulating the at least one modulated radio signal.
7 . The method of claim 1 , further comprising:
transmitting, from a speaker device to the voice-controlled device, the electromagnetic signal; and producing, by the speaker device, at least some of the other sound using the second audio signal.
8 . The method of claim 7 , further comprising generating, at the speaker device, the electromagnetic signal by modulating a radio signal using the second audio signal.
9 . The method of claim 7 , further comprising:
receiving, at the speaker device, an electrical signal through one or more conductors, the electrical signal comprising the second audio signal; and powering a transmitter for the electromagnetic signal in the speaker device using the electrical signal.
10 . The method of claim 1 , wherein the electromagnetic signal comprises a radio-frequency carrier modulated using the second audio signal.
11 . The method of claim 10 , wherein the radio-frequency carrier has a frequency of less than 300 MHz.
12 . The method of claim 1 , wherein the electromagnetic signal comprises one or more electromagnetic signals, the method further comprising:
obtaining a third audio signal from the one or more electromagnetic signals; and correlating the first audio signal with the third audio signal to calculate a correlation value; in response to the correlation value being larger than a threshold, further reducing the contribution from the other sound in the first audio signal by using the third audio signal to generate the processed audio signal.
13 . The method of claim 1 , further comprising:
determining that both a first version of the second audio signal and a second version of the second audio signal are present within the first audio signal, wherein the first version of the second audio signal and the second version of the second audio signal each have at least one of a different amplitude than, a delay from, or a frequency shift from, the second audio signal and from each other; and subtracting both the first version of the second audio signal and the second version of the second audio signal from the first audio signal as at least a part of said processing.
14 . The method of claim 1 , further comprising:
receiving at least some of the other sound at a transducer of a second device remote from the voice-controlled device; converting the received at least some of the other sound into the second audio signal; generating the electromagnetic signal using the second audio signal; and transmitting the electromagnetic signal from the second device for reception by the voice-controlled device.
15 . The method of claim 1 , wherein the electromagnetic signal comprises a wireless radio signal.
16 . The method of claim 1 , wherein the electromagnetic signal is transmitted through a wired network medium.
17 . The method of claim 1 , further comprising:
detecting one or more versions of the second audio signal within the first audio signal; determining an acoustic transfer function that maps the second audio signal to the detected one or more versions of the second audio signal; and using the determined acoustic transfer function to remove the one or more versions of the second audio signal from the first audio signal as at least a part of said processing.
18 . The method of claim 1 , further comprising:
determining identifying information for the received electromagnetic signal; retrieving one or more previously stored characteristics based on the identifying information; and using the retrieved characteristics with the second audio signal as at least a part of said processing.
19 . A voice-controlled device comprising:
a microphone configured to receive a set of sound waves comprising speech uttered by a user and other sound, and to output a first audio signal that includes a contribution from the speech uttered by the user and a contribution from the other sound; a receiver configured to receive an electromagnetic signal and to output a second audio signal obtained from the electromagnetic signal; and an audio pre-processor configured to process the first audio signal using the second audio signal to reduce the contribution from the other sound in a processed audio signal; wherein the voice-controlled device is configured to provide the processed audio signal to a speech recognition module to determine a voice command issued by the user.
20 . The voice-controlled device of claim 19 , further comprising:
a correlator configured to correlate the first audio signal with the second audio signal and to generate one or more correlation parameters; wherein the audio pre-processor is configured to reduce the contribution from the other sound in the processed audio signal by using the one or more correlation parameters with the second audio signal.
21 . The voice-controlled device of claim 19 , further comprising:
a cross-correlator configured to receive the first audio signal and the second audio signal and apply a cross correlation function to provide an output to the audio preprocessor; wherein the audio pre-processor is further configured to determine a time delay and/or a scaling factor based on the output of the cross-correlator, and to use the time delay and/or the scaling factor with the second audio signal to reduce the contribution from the other sound in the processed audio signal.
22 . The voice-controlled device of claim 19 , wherein the receiver is further configured to receive one or more electromagnetic signals to obtain a plurality of other audio signals, including the second audio signal, and wherein the audio pre-processor is further configured to use at least one of the plurality of other audio signals, in addition to the second audio signal, to reduce the contribution from the other sound in the processed audio signal.
23 . The voice-controlled device of claim 19 , wherein the electromagnetic wave comprises a wireless radio signal that has a frequency less than 300 MHz.
24 . The voice-controlled device of claim 19 , wherein the voice-controlled device is configured to determine signal characteristics for a plurality of copies of the second audio signal that are present within the first audio signal and wherein the audio pre-processor is further configured to process the first audio signal based on the signal characteristics to generate the processed audio signal.
25 . The voice-controlled device of claim 19 , wherein the electromagnetic signal is received from a remote device and the second audio signal represents sounds that are captured by a transducer at the remote device.
26 . The voice-controlled device of claim 19 , wherein the voice-controlled device is configured to:
receive, at the receiver, a plurality of electromagnetic signals, including the electromagnetic signal, from a plurality of speaker devices, and output a plurality of other audio signals, including the second audio signal, obtained from the plurality of electromagnetic signals; receive, at the microphone, sound waves from the plurality of speaker devices as at least a part of the other sound; and wherein the audio pre-processor is further configured to use at least one of the plurality of other audio signals, in addition to the second audio signal, to reduce the contribution from the other sound in the processed audio signal.
27 . The voice-controlled device of claim 19 , wherein the audio pre-processor is further configured to:
detect one or more versions of the second audio signal within the first audio signal; determine an acoustic transfer function that maps the second audio signal to the detected one or more versions of the second audio signal; and use the determined acoustic transfer function to remove the one or more versions of the second audio signal from the first audio signal to generate the processed audio signal.
28 . The voice-controlled device of claim 19 , further comprising an antenna coupled to the receiver, the antenna configured to wirelessly receive the electromagnetic signal, the electromagnetic signal comprising a radio-frequency carrier modulated using the second audio signal.
29 . The voice-controlled device of claim 19 , further comprising a connector coupled to the receiver, the connector configured to receive the electromagnetic signal over one or more electrical conductors.
30 . The voice-controlled device of claim 19 , wherein the receiver is further configured to determine identifying information for the received electromagnetic signal; and
the audio pre-processor is further configured to retrieve one or more previously stored characteristics based on the identifying information and use the retrieved characteristics with the second audio signal to reduce the contribution from the other sound in the processed audio signal.
31 . A non-transitory computer-readable storage medium storing instructions which, when executed by at least one processor, program the at least one processor to:
obtain a first audio signal that includes a contribution from speech uttered by a user and a contribution from other sound, the first audio signal derived from a set of sound waves received at a microphone of a voice-controlled device, the set of sound waves comprising the speech uttered by the user and the other sound; obtain a second audio signal, the second audio signal derived from an electromagnetic signal received at a receiver of the voice-controlled device; correlate the first audio signal and the second audio signal to generate one or more correlation parameters, the one or more correlation parameters indicating a time delay and/or a scaling factor for the second audio signal; reduce the contribution from the other sound in the first audio signal using the one or more correlation parameters to generate a processed audio signal; and provide the processed audio signal to a speech recognition module to determine a voice command issued by the user.
32 . The storage medium of claim 31 , the at least one processor further programmed to:
obtain a plurality of other audio signals, including the second audio signal, from the electromagnetic signal, the electromagnetic signal comprising one or more electromagnetic signals; detect one or more of the plurality of other audio signals within the first audio signal; and process the first audio signal using the detected one or more of the plurality of other audio signals to reduce the contribution from the other sound in the processed audio signal.
33 . The storage medium of claim 32 , wherein the one or more electromagnetic signals comprise at least one modulated radio signal and the plurality of other audio signals are obtained by demodulating the at least one modulated radio signal.
34 . The storage medium of claim 31 , the at least one processor further programmed to:
obtain a third audio signal from the electromagnetic signal, the electromagnetic signal comprising one or more electromagnetic signals; correlate the first audio signal with the third audio signal to calculate a correlation value; and in response to the correlation value being larger than a threshold, further reduce the contribution from the other sound in the first audio signal by using the third audio signal to generate the processed audio signal.
35 . The storage medium of claim 31 , wherein the one or more correlation parameters indicate a plurality of time delays and/or scaling factors for the second audio signal due to a plurality of versions of the second audio signal being found in the first audio signal.
36 . The storage medium of claim 31 , the at least one processor further programmed to:
determine identifying information for the received electromagnetic signal; retrieve one or more previously stored characteristics based on the identifying information; and use the retrieved characteristics with the second audio signal as at least a part of said processing.Join the waitlist — get patent alerts
Track US2021312920A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.