Acoustic echo cancellation system and associated method
Abstract
An acoustic echo cancellation (AEC) system includes an adaptive filter, a subtraction circuit, and a processor executing a model. The adaptive filter is arranged to generate an estimated echo signal according to a first microphone signal played by a loudspeaker. The subtraction circuit is arranged to subtract the estimated echo signal from a signal that is output from a microphone receiving both of a speech signal and an echo signal, to generate a second microphone signal, wherein the first microphone signal is not output from the microphone, and the echo signal is transmitted from the loudspeaker to the microphone. The model is arranged to perform short-time Fourier transform upon the first microphone signal and the second microphone signal, respectively, and generate an estimated speech signal through a neural network according to a first transformed microphone signal and a second transformed microphone signal.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An acoustic echo cancellation (AEC) system, comprising:
an adaptive filter, arranged to generate an estimated echo signal according to a first microphone signal played by a loudspeaker; a subtraction circuit, arranged to subtract the estimated echo signal from a signal that is output from a microphone receiving both of a speech signal and an echo signal, to generate a second microphone signal, wherein the first microphone signal is not output from the microphone, and the echo signal is transmitted from the loudspeaker to the microphone; and a processor, arranged to execute:
a model, arranged to perform short-time Fourier transform upon the first microphone signal and the second microphone signal, respectively, to generate a first transformed microphone signal and a second transformed microphone signal, and generate an estimated speech signal through a neural network according to the first transformed microphone signal and the second transformed microphone signal.
2 . The AEC system of claim 1 , wherein the model comprises:
multiple segment modules, arranged to split the first microphone signal and the second microphone signal, respectively, to generate a first segmented microphone signal and a second segmented microphone signal; multiple fast Fourier transform modules, arranged to perform fast Fourier transform upon the first segmented microphone signal and the second segmented microphone signal, respectively, to generate the first transformed microphone signal and the second transformed microphone signal; multiple instant layer normalization (iLN) modules, arranged to normalize the first transformed microphone signal and the second transformed microphone signal, respectively, to generate a first normalized microphone signal and a second normalized microphone signal; a concat module, arranged to concatenate the first normalized microphone signal and the second normalized microphone signal, to generate a concatenated result; and a separation kernel, arranged to predict and generate two masks according to the concatenated result, wherein the estimated speech signal is generated according to the two masks.
3 . The AEC system of claim 2 , wherein a noisy speech signal is a sum of a clean speech signal and a noise signal; the two masks are a real part mask and an imaginary part mask, the real part mask and the imaginary part mask are a real part and an imaginary part of a ratio of a spectral magnitude of the clean speech signal and a spectral magnitude of the noisy speech signal, respectively; the real part mask corresponds to magnitude information of the first microphone signal and the second microphone signal, and the imaginary part mask corresponds to phase information of the first microphone signal and the second microphone signal.
4 . The AEC system of claim 3 , wherein a real part of the estimated speech signal is obtained by multiplying the real part mask by a real part of the second microphone signal, and an imaginary part of the estimated speech signal is obtained by multiplying the imaginary part mask by an imaginary part of the second microphone signal.
5 . The AEC system of claim. 2 , wherein the separation kernel comprises multiple long short term memory (LSTM) layers and a fully-connected layer with sigmoid activation, and the two masks are predicted and generated by the multiple LSTM layers and the fully-connected layer with sigmoid activation.
6 . The AEC system of claim 1 , wherein the adaptive filter is arranged to multiply a room response and an impulse response between the loudspeaker and the microphone, to generate the estimated echo signal.
7 . An acoustic echo cancellation (AEC) method, comprising:
generating an estimated echo signal according to a first microphone signal played by a loudspeaker; subtracting the estimated echo signal from a signal that is output from a microphone receiving both of a speech signal and an echo signal, to generate a second microphone signal, wherein the first microphone signal is not output from the microphone, and the echo signal is transmitted from the loudspeaker to the microphone; performing short-time Fourier transform upon the first microphone signal and the second microphone signal, respectively, to generate a first transformed microphone signal and a second transformed microphone signal; and generating an estimated speech signal through a neural network according to the first transformed microphone signal and the second transformed microphone signal.
8 . The AEC method of claim 7 , wherein the step of performing the short-time Fourier transform upon the first microphone signal and the second microphone signal, respectively, to generate the first transformed microphone signal and the second transformed microphone signal comprises:
splitting the first microphone signal and the second microphone signal, respectively, to generate a first segmented microphone signal and a second segmented microphone signal; and performing fast Fourier transform upon the first segmented microphone signal and the second segmented microphone signal, respectively, to generate the first transformed microphone signal and the second transformed microphone signal.
9 . The AEC method of claim 7 , wherein the step of generating the estimated speech signal through the neural network according to the first transformed microphone signal and the second transformed microphone signal comprises:
normalizing the first transformed microphone signal and the second transformed microphone signal, respectively, to generate a first normalized microphone signal and a second normalized microphone signal; concatenating the first normalized microphone signal and the second normalized microphone signal, to generate a concatenated result; and predicting and generating two masks according to the concatenated result, wherein the estimated speech signal is generated according to the two masks.
10 . The AEC method of claim 9 , wherein a noisy speech signal is a sum of a clean speech signal and a noise signal; the two masks are a real part mask and an imaginary part mask, the real part mask and the imaginary part mask are a real part and an imaginary part of a ratio of a spectral magnitude of the clean speech signal and a spectral magnitude of the noisy speech signal, respectively; the real part mask corresponds to magnitude information of the first microphone signal and the second microphone signal, and the imaginary part mask corresponds to phase information of the first microphone signal and the second microphone signal.
11 . The AEC method of claim 10 , wherein the method further comprises:
obtaining a real part of the estimated speech signal by multiplying the real part mask by a real part of the second microphone signal; and obtaining an imaginary part of the estimated speech signal by multiplying the imaginary part mask by an imaginary part the second microphone signal.
12 . The AEC method of claim 9 , wherein the step of predicting and generating the two masks according to the concatenated result comprises:
predicting and generating the two masks by multiple long short term memory (LSTM) layers and a fully-connected layer with sigmoid activation.
13 . The AEC method of claim 7 , wherein the step of generating the estimated echo signal according to the first microphone signal comprises:
multiplying a room response and an impulse response between the loudspeaker and the microphone, to generate the estimated echo signal.Join the waitlist — get patent alerts
Track US2024048905A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.