US2024371389A1PendingUtilityA1

Neural noise reduction with linear and nonlinear filtering for single-channel audio signals

Assignee: SYNAPTICS INCPriority: May 2, 2023Filed: May 2, 2023Published: Nov 7, 2024
Est. expiryMay 2, 2043(~16.8 yrs left)· nominal 20-yr term from priority
G10L 25/30G10L 21/0364G10L 21/034G10L 25/84G10L 21/0232G10L 25/06G10L 25/18G10L 21/0264
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques that combine statistical signal processing with neural network inferencing. In some aspects, a speech enhancement system may include a linear filter, a deep neural network (DNN), and a nonlinear post-filter. The linear filter and the nonlinear post-filter are configured to suppress noise in audio signals using statistical signal processing techniques. More specifically, the linear filter denoises an input audio signal based on a temporal correlation between successive frames of the audio signal. The DNN infers a speech signal and a noise signal (representing a speech component and a noise component, respectively, of the audio signal) based on the denoised audio signal. The nonlinear post-filter suppresses residual noise in the speech signal based on one or more Gaussian mixture models (GMM) associated with the speech signal and the noise signal.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of speech enhancement, comprising:
 receiving a series of frames of an audio signal;   denoising a first frame in the series of frames based at least in part on a temporal correlation between the series of frames;   inferring a probability of speech associated with the denoised first frame based on a neural network model;   generating a first speech signal and a first noise signal based on the probability of speech associated with the denoised first frame, the first speech signal and the first noise signal representing a speech component and a noise component, respectively, of the audio signal in the first frame;   determining a first spectral suppression gain based on the first speech signal and the first noise signal; and   suppressing residual noise in the first speech signal based on the first spectral suppression gain.   
     
     
         2 . The method of  claim 1 , wherein the audio signal comprises a single channel of audio data. 
     
     
         3 . The method of  claim 1 , wherein the first frame is denoised based on a multi-frame minimum variance distortionless response (MF-MVDR) beamformer that reduces a power of the noise component of the audio signal without distorting the speech component. 
     
     
         4 . The method of  claim 1 , further comprising:
 determining an interframe correlation (IFC) vector associated with a speech component of the audio signal based at least in part on the probability of speech associated with the denoised first frame;   denoising a second frame in the series of frames based at least in part on the IFC vector;   inferring a probability of speech associated with the denoised second frame based on the neural network model;   generating a second speech signal and a second noise signal based on the probability of speech associated with the denoised second frame, the second speech signal and the second noise signal representing a speech component and a noise component, respectively, of the audio signal in the second frame;   determining a second spectral suppression gain based on the second speech signal and the second noise signal; and   suppressing residual noise in the second speech signal based on the second spectral suppression gain.   
     
     
         5 . The method of  claim 1 , wherein the speech component of the audio signal in the denoised first frame is equal to the speech component of the audio signal in the first speech signal. 
     
     
         6 . The method of  claim 1 , wherein the determining of the first spectral suppression gain comprises:
 determining a number (M) of voice activity detection (VAD) features that are indicative of whether speech is present in the first frame based at least in part on the first speech signal and the first noise signal;   determining M probabilities of speech associated with the first speech signal based on the M VAD features, respectively; and   determining a magnitude or power of the residual noise in the first speech signal based on the M probabilities of speech associated with the first speech signal.   
     
     
         7 . The method of  claim 6 , wherein the magnitude or power of the residual noise in the first speech signal is determined based only on the lowest probability of speech among the M probabilities of speech associated with the first speech signal. 
     
     
         8 . The method of  claim 6 , wherein each of the M probabilities of speech associated with the first speech signal is determined based on a respective Gaussian mixture model (GMM). 
     
     
         9 . The method of  claim 6 , wherein M>1. 
     
     
         10 . The method of  claim 6 , wherein the M VAD features include a normalized difference between the first speech signal and the first noise signal. 
     
     
         11 . The method of  claim 6 , wherein the M VAD features include at least one of a cepstral peak, a spectral entropy, or a harmonic product spectrum (HPS) associated with the first speech signal. 
     
     
         12 . A speech enhancement system comprising:
 a processing system; and   a memory storing instructions that, when executed by the processing system, causes the speech enhancement system to:
 receive a series of frames of an audio signal; 
 denoise a first frame in the series of frames based at least in part on a temporal correlation between the series of frames; 
 infer a probability of speech associated with the denoised first frame based on a neural network model; 
 generate a first speech signal and a first noise signal based on the probability of speech associated with the denoised first frame, the first speech signal and the first noise signal representing a speech component and a noise component, respectively, of the audio signal in the first frame; 
 determine a first spectral suppression gain based on the first speech signal and the first noise signal; and 
 suppress residual noise in the first speech signal based on the first spectral suppression gain. 
   
     
     
         13 . The speech enhancement system of  claim 12 , wherein the audio signal comprises a single channel of audio data. 
     
     
         14 . The speech enhancement system of  claim 12 , wherein the first frame is denoised based on a multi-frame minimum variance distortionless response (MF-MVDR) beamformer that reduces a power of the noise component of the audio signal without distorting the speech component. 
     
     
         15 . The speech enhancement system of  claim 12 , wherein execution of the instructions further causes the speech enhancement system to:
 determine an interframe correlation (IFC) vector associated with a speech component of the audio signal based at least in part on the probability of speech associated with the denoised first frame;   denoise a second frame in the series of frames based at least in part on the IFC vector;   infer a probability of speech associated with the denoised second frame based on the neural network model;   generate a second speech signal and a second noise signal based on the probability of speech associated with the denoised second frame, the second speech signal and the second noise signal representing a speech component and a noise component, respectively, of the audio signal in the second frame;   determine a second spectral suppression gain based on the second speech signal and the second noise signal; and   suppress residual noise in the second speech signal based on the second spectral suppression gain.   
     
     
         16 . The speech enhancement system of  claim 12 , wherein the speech component of the audio signal in denoised first frame is equal to the speech component of the audio signal in the first speech signal. 
     
     
         17 . The speech enhancement system of  claim 12 , wherein the determining of the first spectral suppression gain comprises:
 determining a number (M) of voice activity detection (VAD) features that are indicative of whether speech is present in the first frame based at least in part on the first speech signal and the first noise signal;   determining M probabilities of speech associated with the first speech signal based on the M VAD features, respectively; and   determining a magnitude or power of the residual noise in the first speech signal based on the M probabilities of speech associated with the first speech signal.   
     
     
         18 . The speech enhancement system of  claim 17 , wherein each of the M probabilities of speech associated with the first speech signal is determined based on a respective Gaussian mixture model (GMM). 
     
     
         19 . The speech enhancement system of  claim 17 , wherein M>1. 
     
     
         20 . The speech enhancement system of  claim 17 , wherein the M VAD features include at least one of a normalized difference between the first speech signal and the first noise signal, a cepstral peak of the first speech signal, a spectral entropy of the first speech signal, or a harmonic product spectrum (HPS) associated with the first speech signal.

Join the waitlist — get patent alerts

Track US2024371389A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.