Electronic device and method of low latency speech enhancement using autoregressive conditioning-based neural network model
Abstract
A neural method model is trained by, in an initial training iteration, training the neural network model in a teacher forcing mode in which an autoregressive channel includes a ground-truth shifted waveform, and outputting predictions of the neural network model; and in at least one additional training iteration, replacing the ground-truth shifted waveform in the autoregressive channel with the predictions of the neural network model obtained in a previous training iteration. An inference may then be performed by providing, for the neural network model, an additional channel containing at least one prediction of the neural network model outputted during training; and performing speech enhancement using the neural network model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of training and operating a neural network model, the method comprising, by at least one processor of an electronic device:
in an initial training iteration, training the neural network model in a teacher forcing mode in which an autoregressive channel includes a ground-truth shifted waveform, and outputting predictions of the neural network model; in at least one additional training iteration, replacing the ground-truth shifted waveform in the autoregressive channel with the predictions of the neural network model obtained in a previous training iteration.
2 . The method of claim 1 , wherein the at least one additional training iteration comprises a plurality of training iterations, each training iteration outputting respective predictions of the neural network model to the autoregressive channel for a next iteration of the plurality of training iterations.
3 . The method of claim 2 , wherein the neural network model is configured to perform at least one forward pass, compute a loss, and perform at least one backward pass, and
wherein, during training, a number of forward passes performed before computing the loss and performing the at least one backward pass is gradually increased.
4 . The method of claim 2 , wherein only an output of a final iteration of the plurality of training iterations is backpropagated.
5 . The method of claim 1 , further comprising performing an inference by the neural network model by:
providing, for the neural network model, an additional channel containing at least one prediction of the neural network model outputted during training; and performing speech enhancement using the neural network model.
6 . The method of claim 5 , wherein the neural network model includes a fully convolutional neural network.
7 . The method of claim 6 , wherein the fully convolutional neural network includes a WaveUNet architecture augmented at a bottleneck thereof with a long short term memory (LSTM) layer.
8 . An electronic device comprising:
at least one memory storing at least one instruction; and at least one processor configured to execute the at least one instruction to:
in an initial training iteration, train the neural network model in a teacher forcing mode in which an autoregressive channel includes a ground-truth shifted waveform, and output predictions of the neural network model, and
in at least one additional training iteration, replace the ground-truth shifted waveform in the autoregressive channel with the predictions of the neural network model obtained a previous training iteration.
9 . The electronic device of claim 8 , wherein the at least one additional training iteration comprises a plurality of training iterations, each training iteration outputting respective predictions of the neural network model to the autoregressive channel for a next iteration of the plurality of training iterations.
10 . The electronic device of claim 9 , wherein the neural network model is configured to perform at least one forward pass, compute a loss, and perform at least one backward pass, and
wherein, during training, a number of forward passes performed before computing the loss and performing the at least one backward pass is gradually increased.
11 . The electronic device of claim 9 , wherein only an output of a final iteration of the plurality of training iterations is backpropagated.
12 . The electronic device of claim 8 , wherein the at least one processor is further configured to perform an inference by:
providing, for the neural network model, an additional channel containing at least one prediction of the neural network model outputted during training, and performing speech enhancement using the neural network model.
13 . The electronic device of claim 12 , wherein the neural network model includes a fully convolutional neural network.
14 . The electronic device of claim 13 , wherein the fully convolutional neural network includes a WaveUNet architecture augmented at a bottleneck thereof with a long short term memory (LSTM) layer.
15 . A non-transitory computer-readable medium having instructions stored thereon which, when executed by at least one processor, cause the at least one processor to:
in an initial training iteration, train the neural network model in a teacher forcing mode in which an autoregressive channel includes a ground-truth shifted waveform, and output predictions of the neural network model; and in at least one additional training iteration, replace the ground-truth shifted waveform in the autoregressive channel with the predictions of the neural network model obtained a previous training iteration.
16 . The medium of claim 15 , wherein the at least one additional training iteration comprises a plurality of training iterations, each training iteration outputting respective predictions of the neural network model to the autoregressive channel for a next iteration of the plurality of training iterations.
17 . The medium of claim 16 , wherein the neural network model is configured to perform at least one forward pass, compute a loss, and perform at least one backward pass, and
wherein, during training, a number of forward passes performed before computing the loss and performing the at least one backward pass is gradually increased.
18 . The medium of claim 16 , wherein the instructions further cause the at least one processor to perform an inference by:
providing, for the neural network model, an additional channel containing at least one prediction of the neural network model outputted during training, and performing speech enhancement using the neural network model.
19 . The electronic device of claim 15 , wherein the neural network model includes a fully convolutional neural network.
20 . The electronic device of claim 19 , wherein the fully convolutional neural network includes a WaveUNet architecture augmented at a bottleneck thereof with a long short term memory (LSTM) layer.Join the waitlist — get patent alerts
Track US2024161736A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.