US2016322067A1PendingUtilityA1
Methods and Voice Activity Detectors for a Speech Encoders
Assignee: ERICSSON TELEFON AB L M (publ)Priority: Oct 19, 2009Filed: Jun 14, 2016Published: Nov 3, 2016
Est. expiryOct 19, 2029(~3.2 yrs left)· nominal 20-yr term from priority
Inventors:Martin Sehlstedt
G10L 25/51G10L 21/0208G10L 2025/786G10L 25/18G10L 25/78G10L 25/87
46
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Voice activity detectors and related methods are provided. Methods include receiving a frame of the input signal; determining a first SNR of the received frame; comparing the determined first SNR with an adaptive threshold; and detecting whether the received frame comprises voice based on the comparison. The adaptive threshold is at least based on total noise energy of a noise level, an estimate of a second SNR and on energy variation between different frames.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A computer-implemented method, in a voice activity detector, for determining whether frames of an input signal comprise voice comprising:
receiving a frame of the input signal; determining a first signal-to-noise-ratio (SNR) of the received frame by:
obtaining a plurality of subband SNR values of the received frame by dividing energy levels of each of a plurality of subbands of the received frame by its respective subband background energy;
applying subband specific significance thresholds to each of the plurality of subband SNR values of the received frame through selectively adjusting the plurality of subband SNR values using a non-linear function; and
obtaining the first SNR by summing together all non-linearly adjusted SNR values of each of the plurality of subbands;
comparing the determined first SNR with an adaptive threshold,
wherein the adaptive threshold is at least based on an estimate of a second SNR and an estimate of energy variation between different frames of the input signal,
wherein the second SNR is a long-term SNR, and
wherein the estimate of energy variation between different frames has limits on how quickly the estimate increases such that the estimate may not increase beyond a fixed constant for each frame; and
detecting whether the received frame comprises voice based on the comparison.
3 . The method of claim 2 , wherein the energy variation between different frames is the energy variation between the received frame and a last received frame which did not comprise voice.
4 . The method of claim 2 , wherein the second SNR is measured over a plurality of frames of the input signal.
5 . The method of claim 2 wherein comparing the determined first SNR with the adaptive threshold comprises adjusting the estimate of the second SNR upwards responsive to a determination that the estimate of the second SNR is lower than a smooth input dynamics measure, wherein the smooth input dynamics measure is indicative of energy dynamics of the received frame.
6 . The method of claim 5 , wherein the estimate of the second SNR is adjusted upwards to a value which is less than or equal to the smooth input dynamics measure.
7 . The method of claim 6 , wherein the smooth input dynamics measure for the received frame comprises a high/max energy tracker, a low/min energy tracker, and a portion corresponding to a smooth input dynamics measure for prior frames.
8 . The method of claim 7 , wherein, prior to calculating the smooth input dynamics measure for the received frame, the low/min energy tracker is set to either the total energy of the received frame or an incremental upward adjustment of the low/min energy tracker, whichever is smaller, and
wherein, prior to calculating the smooth input dynamics measure for the received frame, the high/max energy tracker is set to either the total energy of the received frame or a decremental downward adjustment of the high/max energy tracker, whichever is greater.
9 . The method of claim 7 , wherein the smooth input dynamics measure is a function of a difference between the high/Max energy tracker based on a highest frame energy value over a plurality of frames and the low/mm n energy tracker based on a lowest frame energy value over the plurality of frames.
10 . A voice activity detector for determining whether frames of an input signal comprise voice, the voice activity detector comprising:
an input circuit configured to receive a frame of the input signal; and a processor configured to:
determine a first signal-to-noise-ratio (SNR) of the received frame by:
obtaining a plurality of subband SNR values of the received frame by dividing energy levels of each of a plurality of subbands of the received frame by its respective subband background energy;
applying subband specific significance thresholds to each of the plurality of subband SNR values of the received frame through selectively adjusting the plurality of subband SNR values using a non-linear function; and
obtaining the first SNR by summing together all non-linearly adjusted SNR values of each of the plurality of subbands;
compare the determined first SNR with an adaptive threshold,
wherein the adaptive threshold is at least based on an estimate of a second SNR and an estimate of energy variation between different frames of the input signal,
wherein the second SNR is a long-term SNR, and
wherein the estimate of energy variation between different frames has limits on how quickly the estimate increases such that the estimate may not increase beyond a fixed constant for each frame; and
detect whether the received frame comprises voice based on the comparison.
11 . The voice activity detector of claim 10 , wherein the energy variation between different frames is the energy variation between the received frame and a last received frame which did not comprise voice.
12 . The voice activity detector of claim 10 , wherein the second SNR is measured over a plurality of frames of the input signal.
13 . The voice activity detector of claim 10 wherein comparing the determined first SNR with the adaptive threshold comprises adjusting the estimate of the second SNR upwards responsive to a determination that the estimate of the second SNR is lower than a smooth input dynamics measure, wherein the smooth input dynamics measure is indicative of energy dynamics of the received frame.
14 . The voice activity detector of claim 13 , wherein the estimate of the second SNR is adjusted upwards to a value which is less than or equal to the smooth input dynamics measure.
15 . The voice activity detector of claim 13 , wherein the smooth input dynamics measure for the received frame comprises a high/max energy tracker, a low/min energy tracker, and a portion corresponding to a smooth input dynamics measure for prior frames.
16 . The voice activity detector of claim 15 , wherein, prior to calculating the smooth input dynamics measure for the received frame, the low/mm energy tracker is set to either the total energy of the received frame or an incremental upward adjustment of the low/min energy tracker, whichever is smaller, and
wherein, prior to calculating the smooth input dynamics measure for the received frame, the high/max energy tracker is set to either the total energy of the received frame or a decremental downward adjustment of the high/max energy tracker, whichever is greater.)
17 . The voice activity detector of claim 15 , wherein the smooth input dynamics measure is a function of a difference between the high/max energy tracker based on a highest frame energy value over a plurality of frames and the low/min energy tracker based on a lowest frame energy value over the plurality of frames.Join the waitlist — get patent alerts
Track US2016322067A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.