Audio decoder, audio encoder, and related methods using joint coding of scale parameters for channels of a multi-channel audio signal
Abstract
Audio decoder for decoding an encoded audio signal having multi-channel audio data having data for two or more audio channels, and information on jointly encoded scale parameters, having: a scale parameter decoder for decoding the information on the jointly encoded scale parameters to obtain a first and a second set of scale parameters for a first channel and a second channel, respectively, of a decoded audio signal; and a signal processor for applying the first and second sets of scale parameters to a first and second channel representation, respectively, derived from the multi-channel audio data to obtain the first and second channels of the decoded audio signal, wherein the jointly encoded scale parameters have information on a first group and on a second group of jointly encoded scale parameters, and wherein the scale parameter decoder is configured to combine a jointly encoded scale parameter of the first group and one of the second group using a first and a second combination rule, respectively, to obtain a scale parameter of the first and second sets of scale parameters.
Claims
exact text as granted — not AI-modified1 . An audio decoder for decoding an encoded audio signal comprising multi-channel audio data comprising data for two or more audio channels, and information on jointly encoded scale parameters, comprising:
a scale parameter decoder for decoding the information on the jointly encoded scale parameters to acquire a first set of scale parameters for a first channel of a decoded audio signal and a second set of scale parameters for a second channel of the decoded audio signal; and a signal processor for applying the first set of scale parameters to a first channel representation derived from the multi-channel audio data and for applying the second set of scale parameters to a second channel representation derived from the multi-channel audio data to acquire the first channel and the second channel of the decoded audio signal, wherein the jointly encoded scale parameters comprise information on a first group of jointly encoded scale parameters and information on a second group of jointly encoded scale parameters, and wherein the scale parameter decoder is configured to combine a jointly encoded scale parameter of the first group and a jointly encoded scale parameter of the second group using a first combination rule to acquire a scale parameter of the first set of scale parameters, and using a second combination rule being different from the first combination rule to acquire a scale parameter of the second set of scale parameters.
2 . The audio decoder of claim 1 , wherein the first group of jointly encoded scale parameters comprises mid scale parameters and the second group of jointly encoded scale parameters comprises side scale parameters, and wherein the scale parameter decoder is configured to use, in the first combination rule, an addition, and to use, in the second combination rule, a subtraction.
3 . The audio decoder of claim 1 , wherein the encoded audio signal is organized in a sequence of frames, wherein a first frame comprises the multi-channel audio data and the information on the jointly encoded scale parameters, and wherein a second frame comprises separately encoded scale parameter information, and
wherein the scale parameter decoder is configured to detect that the second frame comprises the separately encoded scale parameter information and to calculate the first set of scale parameters and the second set of scale parameters.
4 . The audio decoder of claim 3 , wherein the first frame and the second frame each comprise a state side information indicating, in a first state, that the first frame comprises the information on the jointly encoded scale parameters and, in a second state, that the second frame comprises the separately encoded scale parameter information, and
wherein the scale parameter decoder is configured to read the state side information of the second frame, to detect that the second frame comprises the separately encoded scale parameter information based on the state side information read, or to read the state side information of the first frame, and to detect that the first frame comprises the information on the jointly encoded scale parameters using the state side information read.
5 . The audio decoder of claim 1 ,
wherein the signal processor is configured to decode the multi-channel audio data to derive the first channel representation and the second channel representation, wherein the first channel representation and the second channel representation are spectral domain representations comprising spectral sampling values, and wherein the signal processor is configured to apply each scale parameter of the first set and the second set to a corresponding plurality of the spectral sampling values to acquire a shaped spectral representation of the first channel and a shaped spectral representation of the second channel.
6 . The audio decoder of claim 5 , wherein the signal processor is configured to convert the shaped spectral representation of the first channel and the shaped spectral representation of the second channel into a time domain to acquire a time domain representation of the first channel and a time domain representation of the second channel of the decoded audio signal.
7 . The audio decoder of claim 1 , wherein the first channel representation comprises a first number of bands, wherein the first set of scale parameters comprises a second number of scale parameters, the second number being lower than the first number, and
wherein the signal processor is configured to interpolate the second number of scale parameters to acquire a number of interpolated scale parameters being greater than or equal to the first number of bands, and wherein the signal processor is configured to scale the first channel representation using the interpolated scale parameters, or wherein the first channel representation comprises a first number of bands, wherein the information on the first group of jointly encoded scale parameters comprises a second number of jointly encoded scale parameters, the second number being lower than the first number, wherein the scale parameter decoder is configured to interpolate the second number of jointly encoded scale parameters to acquire a number of interpolated jointly encoded scale parameters being greater than or equal to the first number of bands, and wherein the scale parameter decoder is configured to process the interpolated jointly encoded scale parameters to determine the first set of scale parameters and the second set of scale parameters.
8 . The audio decoder of claim 1 , wherein the encoded audio signal is organized in a sequence of frames, wherein the information on the second group of jointly encoded scale parameters comprises, in a certain frame, a zero side information, wherein the scale parameter decoder is configured to detect the zero side information to determine that the second group of jointly encoded scale parameters are all zero for the certain frame, and
wherein the scale parameter decoder is configured to derive the scale parameters of the first set of scale parameters and the second set of scale parameters only from the first group of jointly encoded scale parameters or to set, in the combining the jointly encoded scale parameter of the first group and the jointly encoded scale parameter of the second group, to zero values or values being smaller than a noise threshold.
9 . The audio decoder of claim 1 ,
wherein the scale parameter decoder is configured
to de-quantize the information on the first group of jointly encoded scale parameters using a first de-quantization mode, and
to de-quantize the information on the second group of jointly encoded scale parameters using a second de-quantization mode, the second de-quantization mode being different from the first de-quantization mode.
10 . The audio decoder of claim 9 , wherein the scale parameter decoder is configured to use the second de-quantization mode having associated a lower or higher quantization precision than the first de-quantization mode.
11 . The audio decoder of claim 9 , wherein the scale parameter decoder is configured to use, as the first de-quantization mode, a first de-quantization stage and a second de-quantization stage and a combiner, the combiner receiving, as an input, a result of the first de-quantization stage and a result of the second de-quantization stage, and
to use, as the second de-quantization mode, the second de-quantization stage of the first de-quantization mode receiving, as an input, the information on the second group of jointly encoded scale parameters.
12 . The audio decoder of claim 11 , wherein the first de-quantization stage is a vector de-quantization stage and wherein the second de-quantization stage is an algebraic vector de-quantization stage, or wherein the first de-quantization stage is a fixed rate de-quantization stage and wherein the second de-quantization stage is a variable rate de-quantization stage.
13 . The audio decoder of claim 11 , wherein the information on the first group of jointly encoded scale parameters comprises, for a frame of the encoded audio signal, two or more indexes and wherein the information on the second group of jointly encoded scale parameters comprises a single index or a lower number of indexes or the same number of indexes as in the first group, and
wherein the scale parameter decoder is configured to determine, in the first de-quantization stage e.g., for each index of the two or more indexes, intermediate jointly encoded scale parameters of the first group, and wherein the scale parameter decoder is configured to calculate, in the second de-quantization stage, residual jointly encoded scale parameters of the first group e.g. from the single or lower or the same number of indexes of the information on the first group of jointly encoded scale parameters and to calculate, by the combiner the first group of jointly encoded scale parameters from the intermediate jointly encoded scale parameters of the first group and the residual jointly encoded scale parameters of the first group.
14 . The audio decoder of claim 11 , wherein the first de-quantization stage comprises using an index for a first codebook comprising a first number of entries or using an index representing a first precision, wherein the second de-quantization stage comprises using an index for a second codebook comprising a second number of entries or using an index representing a second precision, and wherein the second number is lower or higher than the first number or the second precision is lower or higher than the first precision.
15 . The audio decoder of claim 1 , wherein the information on the second group of jointly encoded scale parameters indicates that the second group of jointly encoded scale parameters are all zero or at a certain value for a frame of the encoded audio signal, and wherein the scale parameter decoder is configured to use, in the combining using the first rule or the second rule, a jointly encoded scale parameter being zero or being at the certain value or being a synthesized jointly encoded scale parameter, or
wherein, for the frame comprising the all zero or certain value information, the scale parameter decoder is configured to determine the second set of scale parameters only using the first group of jointly encoded scale parameters without a combining operation.
16 . The audio decoder of claim 9 , wherein the scale parameter decoder is configured to use, as the first de-quantization mode, the first de-quantization stage and the second de-quantization stage and the combiner, the combiner receiving, as an input, a result of the first de-quantization stage and a result of the second de-quantization stage, and to use, as the second de-quantization smoke, the first de-quantization stage of the first de-quantization mode.
17 . An audio encoder for encoding a multi-channel audio signal comprising two or more channels, comprising:
a scale parameter calculator for calculating a first group of jointly encoded scale parameters and a second group of jointly encoded scale parameters from a first set of scale parameters for a first channel of the multi-channel audio signal and from a second set of scale parameters for a second channel of the multi-channel audio signal; a signal processor for applying the first set of scale parameters to the first channel of the multi-channel audio signal and for applying the second set of scale parameters to the second channel of the multi-channel audio signal and for deriving multi-channel audio data; and an encoded signal former for using the multi-channel audio data and information on the first group of jointly encoded scale parameters and information on the second group of jointly encoded scale parameters to acquire an encoded multi-channel audio signal.
18 . The audio encoder of claim 17 , wherein the signal processor is configured, in the applying,
to encode the first group of jointly encoded scale parameters and the second group of jointly encoded scale parameters to acquire the information on the first group of jointly encoded scale parameters and the information on the second group of jointly encoded scale parameters, to locally decode the information on the first and the second groups of jointly encoded scale parameters to acquire a locally decoded first set of scale parameters and a locally decoded second set of scale parameters, and to scale the first channel using the locally decoded first set of scale parameters and to scale the second channel using the locally decoded second set of scale parameters, or wherein the signal processor is configured, in the applying, to quantize the first group of jointly encoded scale parameters and the second group of jointly encoded scale parameters to acquire a quantized first group of jointly encoded scale parameters and a quantized second group of jointly encoded scale parameters, to locally decode the quantized first and the second groups of jointly encoded scale parameters to acquire a locally decoded first set of scale parameters and a locally decoded second set of scale parameters, and to scale the first channel using the locally decoded first set of scale parameters and to scale the second channel using the locally decoded second set of scale parameters.
19 . The audio encoder of claim 17 ,
wherein the scale parameter calculator is configured to combine a scale parameter of the first set of scale parameters and a scale parameter of the second set of scale parameters using a first combination rule to acquire a jointly encoded scale parameter of the first group of jointly encoded scale parameters, and using a second combination rule different from the first combination rule to acquire a jointly encoded scale parameter of the second group of jointly encoded scale parameters.
20 . The audio encoder of claim 19 , wherein the first group of jointly encoded scale parameters comprises mid scale parameters and the second group of jointly encoded scale parameters comprises side scale parameters, and wherein the scale parameter calculator is configured to use, in the first combination rule, an addition, and to use, in the second combination rule, a subtraction.
21 . The audio encoder of claim 17 , wherein the scale parameters calculator is configured to process a sequence of frames of the multi-channel audio signal,
wherein the scale parameter calculator is configured
to calculate first and second groups of jointly encoded scale parameters for a first frame of the sequence of frames, and
to analyze a second frame of the sequence of frames to determine a separate coding mode for the second frame, and
wherein the encoded signal former is configured to introduce a state side information into the encoded audio signal indicating a separate encoding mode for the second frame or a joint encoding mode for the first frame, and information on the first set and the second set of separately encoded scale parameters for the second frame.
22 . The audio encoder of claim 17 , wherein the scale parameter calculator is configured
to calculate the first set of scale parameters for the first channel and the second set of scale parameters for the second channel, to downsample the first and the second sets of scale parameters to acquire a downsampled first set and a downsampled second set; and to combine a scale parameter from the downsampled first set and the downsampled second set using different combination rules to acquire a jointly encoded scale parameter of the first group and a jointly encoded scale parameter of the second group, or
wherein the scale parameter calculator is configured
to calculate the first set of sale parameters for the first channel and the second set of scale parameters for the second channel,
to combine a scale parameter from the first set and a scale parameter from the second set using different combination rules to acquire a jointly encoded scale parameter of the first group and a jointly encoded scale parameter of the second group, and
to downsample the first group of jointly encoded scale parameters to acquire a downsampled first group of jointly encoded scale parameters, and to downsample the second group of jointly encoded scale parameters to acquire a downsampled second group of jointly encoded scale parameters,
wherein the downsampled first group and the downsampled second group represent the information on the first group of jointly encoded scale parameters and the information on the second group of jointly encoded scale parameters.
23 . The audio encoder of claim 21 ,
wherein the scale parameter calculator is configured to calculate a similarity of the first channel and the second channel in the second frame and to determine the separate encoding mode in case a calculated similarity is in a first relation to a threshold or to determine the joint encoding mode in case the calculated similarity is in a different second relation to the threshold.
24 . The audio encoder of claim 23 , wherein the scale parameter calculator is configured
to calculate, for the second frame, a difference between the scale parameter of the first set and the scale parameter of the second set for each band, to process each difference of the second frame so that negative signs are removed to acquire processed differences of the second frame, to combine the processed differences to acquire a similarity measure, to compare the similarity measure to the threshold, and to decide in favor of the separate coding mode, when the similarity measure is greater than the threshold, or to decide in favor of the joint coding mode, when the similarity measure is lower than the threshold.
25 . The audio encoder of claim 17 , wherein the signal processor is configured
to quantize the first group of jointly encoded scale parameters using a first stage quantization function to acquire one or more first quantization indexes as a first stage result and to acquire an intermediate first group of jointly encoded scale parameters, to calculate a residual first group of jointly encoded scale parameters from the first group of jointly encoded scale parameters and the intermediate first group of jointly encoded scale parameters, and to quantize the residual first group of jointly encoded scale parameters using a second stage quantization function to acquire one or more quantization indexes as a second stage result.
26 . The audio encoder of claim 17 ,
wherein the signal processor is configured to quantize the second group of jointly encoded scale parameters using a single stage quantization function to acquire one or more quantization indexes as the single stage result, or wherein the signal processor is configured for quantizing the first group of jointly encoded scale parameters using at least a first stage quantization function and a second stage quantization function, and wherein the signal processor is configured for quantizing the second group of jointly encoded scale parameters using a single stage quantization function, wherein the single stage quantization function is selected from the first stage quantization function and the second stage quantization function.
27 . The audio encoder of claim 21 , wherein the scale parameter calculator is configured
to quantize the first set of scale parameters using a first stage quantization function to acquire one or more first quantization indexes as a first stage result and to acquire an intermediate first set of scale parameters, to calculate a residual first set of scale parameters from the first set of scale parameters and the intermediate first set of scale parameters, and to quantize the residual first set of scale parameters using a second stage quantization function to acquire one or more quantization indexes as a second stage result,
or
wherein the scale parameter calculator is configured
to quantize the second set of scale parameters using a first stage quantization function to acquire one or more first quantization indexes as a first stage result and to acquire an intermediate second set of scale parameters,
to calculate a residual second set of scale parameters from the second set of scale parameters and the intermediate second set of scale parameters, and
to quantize the residual second set of scale parameters using a second stage quantization function to acquire one or more quantization indexes as a second stage result.
28 . The audio encoder of claim 25 ,
wherein the second stage quantization function uses an amplification or weighting value lower than 1 to increase the residual first group of jointly encoded scaling parameters or the residual first or second set of scale parameters before performing a vector quantization, wherein the vector quantization is performed using increased residual values, and/or wherein, exemplarily, the weighting or amplification value is used to divide a scaling parameter by the weighting or amplification value, wherein the weighting value is advantageously between 0.1 and 0.9, or more advantageously between 0.2 and 0.6 or even more advantageously between 0.25 and 0.4, and/or wherein the same amplification value is used for all scaling parameters of the residual first group of jointly encoded scaling parameters or the residual first or second set of scale parameters.
29 . The audio encoder of claim 25 ,
wherein the first stage quantization function comprises at least one codebook with a first number of entries corresponding to a first size of the one or more quantization indexes, wherein the second stage quantization function or the single stage quantization function comprises at least one codebook with a second number of entries corresponding to a second size of the one or more quantization indexes, and wherein the first number is greater or lower than the second number or the first size is greater or lower than the second size, or wherein the wherein the first stage quantization function is a fixed rate quantization function and wherein the second stage quantization function is a variable rate quantization function.
30 . The audio encoder of claim 17 , wherein the scale parameter calculator is configured
to receive a first MDCT representation for the first channel and a second MDCT representation for the second channel, to receive a first MDST representation for the first channel and a second MDST representation for the second channel, to calculate a first power spectrum for the first channel from the first MDCT representation and the first MDST representation and a second power spectrum for the second channel from the second MDCT representation and the second MDST representation, and to calculate the first set of scale parameters for the first channel from the first power spectrum and to calculate the second set of scale parameters for the second channel from the second power spectrum.
31 . The audio encoder of claim 30 ,
wherein the signal processor is configured to scale the first MDCT representation using information derived from the first set of scale parameters, and to scale the second MDCT representation using information derived from the second set of scale parameters.
32 . The audio encoder of claim 17 ,
wherein the signal processor is configured to further process a scaled first channel representation and a scaled second channel representation using a joint multi-channel processing to derive a multi-channel processed representation of the multi-channel audio signal, to optionally further process using a spectral band replication processing or an intelligent gap filling processing or a bandwidth enhancement processing and to quantize and encode a representation of the channels of the multi-channel audio signal to acquire the multi-channel audio data.
33 . The audio encoder of claim 17 , being configured to determine, for a frame of the multi-channel audio signal, the information on the second group of jointly encoded scale parameters as an all zero or all certain value information indicating the same value or a zero value for all jointly encoded scale parameters of the frame and wherein the encoded signal former is configured to use the all zero or all certain value information to acquire the encoded multi-channel audio signal.
34 . The audio encoder of claim 17 , wherein the scale parameter calculator is configured
for calculating the first group of jointly encoded scale parameters and the second group of jointly encoded scale parameters for a first frame, for calculating the first group of jointly encoded scale parameters for a second frame, wherein, in the second frame, the jointly encoded scale parameters are not calculated or encoded, and
wherein the encoded signal former is configured to use a flag as the information on the second group of jointly encoded scale parameters indicating that, in the second frame, any jointly encoded scale parameters of the second group are not comprised in the encoded multichannel audio signal.
35 . A method of decoding an encoded audio signal comprising multi-channel audio data comprising data for two or more audio channels, and information on jointly encoded scale parameters, comprising:
decoding the information on the jointly encoded scale parameters to acquire a first set of scale parameters for a first channel of a decoded audio signal and a second set of scale parameters for a second channel of the decoded audio signal; and applying the first set of scale parameters to a first channel representation derived from the multi-channel audio data and for applying the second set of scale parameters to a second channel representation derived from the multi-channel audio data to acquire the first channel and the second channel of the decoded audio signal, wherein the jointly encoded scale parameters comprise information on a first group of jointly encoded scale parameters and information on a second group of jointly encoded scale parameters, and wherein the decoding comprises combining a jointly encoded scale parameter of the first group and a jointly encoded scale parameter of the second group using a first combination rule to acquire a scale parameter of the first set of scale parameters, and using a second combination rule being different from the first combination rule to acquire a scale parameter of the second set of scale parameters.
36 . A method of encoding a multi-channel audio signal comprising two or more channels, comprising:
calculating a first group of jointly encoded scale parameters and a second group of jointly encoded scale parameters from a first set of scale parameters for a first channel of the multi-channel audio signal and from a second set of scale parameters for a second channel of the multi-channel audio signal; applying the first set of scale parameters to the first channel of the multi-channel audio signal and applying the second set of scale parameters to the second channel of the multi-channel audio signal and for deriving multi-channel audio data; and using the multi-channel audio data and information on the first group of jointly encoded scale parameters and information on the second group of jointly encoded scale parameters to acquire an encoded multi-channel audio signal.
37 . A non-transitory digital storage medium having stored thereon a computer program for performing a method of decoding an encoded audio signal comprising multi-channel audio data comprising data for two or more audio channels, and information on jointly encoded scale parameters, comprising:
decoding the information on the jointly encoded scale parameters to acquire a first set of scale parameters for a first channel of a decoded audio signal and a second set of scale parameters for a second channel of the decoded audio signal; and applying the first set of scale parameters to a first channel representation derived from the multi-channel audio data and for applying the second set of scale parameters to a second channel representation derived from the multi-channel audio data to acquire the first channel and the second channel of the decoded audio signal, wherein the jointly encoded scale parameters comprise information on a first group of jointly encoded scale parameters and information on a second group of jointly encoded scale parameters, and wherein the decoding comprises combining a jointly encoded scale parameter of the first group and a jointly encoded scale parameter of the second group using a first combination rule to acquire a scale parameter of the first set of scale parameters, and using a second combination rule being different from the first combination rule to acquire a scale parameter of the second set of scale parameters, when said computer program is run by a computer.
38 . A non-transitory digital storage medium having stored thereon a computer program for performing a method of encoding a multi-channel audio signal comprising two or more channels, comprising:
calculating a first group of jointly encoded scale parameters and a second group of jointly encoded scale parameters from a first set of scale parameters for a first channel of the multi-channel audio signal and from a second set of scale parameters for a second channel of the multi-channel audio signal; applying the first set of scale parameters to the first channel of the multi-channel audio signal and applying the second set of scale parameters to the second channel of the multi-channel audio signal and for deriving multi-channel audio data; and using the multi-channel audio data and information on the first group of jointly encoded scale parameters and information on the second group of jointly encoded scale parameters to acquire an encoded multi-channel audio signal, when said computer program is run by a computer.Join the waitlist — get patent alerts
Track US2023133513A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.