Layered coding of audio with discrete objects
Abstract
A first layer of data having a first set of Ambisonic audio components can be decoded where the first set of Ambisonic audio components is generated based on ambience and one or more object-based audio signals. A second layer of data is decoded having at least one of the one or more object-based audio signals. One of the object-based audio signals is subtracted from the first set of Ambisonic audio components. The resulting Ambisonic audio components are rendered to generate a first set of audio channels. The one or more object-based audio signals are spatially rendered to generate a second set of audio channels. Other aspects are described and claimed.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for encoding a layered audio bit stream, comprising:
generating a first set of Ambisonic audio components based on ambience and one or more object-based audio signals; encoding into a bit stream, a first layer having the first set of Ambisonic audio components; and encoding into the bit stream, a second layer having at least one of the one or more object-based audio signals.
2 . The method of claim 1 , further comprising encoding, in the second layer, a second set of Ambisonic audio components, each component of the second set of Ambisonic audio components having a higher order than each component of the first set of Ambisonic audio components.
3 . The method of claim 1 , wherein the first set of Ambisonic audio components includes an omnidirectional pattern (W) audio component, a first bi-directional polar pattern audio component aligned in a first direction (X), a second bi-directional polar pattern audio component aligned in a second direction (Y), and a third bi-directional polar pattern audio component aligned in a third direction (Z), and the first layer of the bit stream does not contain other encoded Ambisonic audio components.
4 . The method of claim 1 , wherein the first set of Ambisonic audio components includes an omnidirectional pattern (W) audio component, first order Ambisonic audio components, and second order Ambisonic audio components, and the first layer of the bit stream does not contain other encoded Ambisonic audio components.
5 . The method of claim 1 , wherein generating the first set of Ambisonic audio components includes converting the one or more object-based audio signals to Ambisonic audio components with the object-based audio, and combining a) the Ambisonic audio components with the object-based audio with b) Ambisonic audio components having the ambience.
6 . The method of claim 5 , wherein the combined ambience and object-based Ambisonic audio components are truncated to remove Ambisonic audio components of an order greater than a threshold, and remaining Ambisonic audio components are the first set of Ambisonic audio components that are encoded in the first layer of the bit stream.
7 . The method of claim 1 , wherein the bit stream includes metadata that indicates a) one or more selectable layers that can be encoded and transmitted in the bit stream, and b) Ambisonic audio components or audio objects in each of the selectable layers.
8 . The method of claim 1 , wherein, at a downstream device,
the at least one of the one or more object-based audio signals is subtracted from the first set of Ambisonic audio components and a resulting set of Ambisonic audio components are rendered by an Ambisonic renderer, into a first set of playback channels; the at least one of the one or more object-based audio signals is spatially rendered into a second set of playback channels; and the first set of playback channels and second set of playback channels are combined into a plurality of speaker channels that are used to drive a plurality of speakers.
9 . The method of claim 1 , wherein the first set of Ambisonic audio components includes a first sublayer having only omnidirectional pattern (W) component.
10 . The method of claim 9 , wherein the first set of Ambisonic audio components includes a second sublayer, having
a) a summation of three Ambisonic audio components, including a1) a first bi-directional polar pattern audio component aligned in a first direction, a2) a second bi-directional polar pattern audio component aligned in a second direction, and a3) a third bi-directional polar pattern audio component aligned in a third direction; and b) one or more parameters generated based on the three Ambisonic audio components.
11 . The method of claim 10 , wherein the first set of Ambisonic audio components includes a third layer that has two of the three Ambisonic audio components that are summed in the second sublayer.
12 . The method of claim 11 , wherein the one or more parameters includes correlation between the three Ambisonic audio components, level differences between the three Ambisonic audio components, or phase differences between the three Ambisonic audio components that are summed in the second sublayer.
13 . The method of claim 12 , wherein weighting coefficients are be applied to each of the three Ambisonic audio components to optimize summation of the three components.
14 . The method of claim 1 , wherein the second layer is encoded in the bit stream only if a bandwidth satisfies a threshold, or if a request from a downstream device is received, and a determination of whether or not to encode the second layer in the bit stream can change from one audio frame to another audio frame.
15 . The method of claim 1 , wherein additional layers, each having a respective set of Ambisonic audio components, are encoded in the bit stream, each additional layer having a same or higher order of Ambisonic audio components than a previous layer.
16 . An audio system, comprising:
a processor and non-transitory computer-readable memory having stored therein instructions that when executed by the processor cause the processor to:
generate a first set of Ambisonic audio components based on ambience and one or more object-based audio signals;
encode into a bit stream, a first layer having the first set of Ambisonic audio components; and
encode into the bit stream, a second layer having at least one of the one or more object-based audio signals.
17 . The audio system of claim 16 , wherein the processor is further to: encode, in the second layer, a second set of Ambisonic audio components, each component of the second set of Ambisonic audio components having a higher order than each component of the first set of Ambisonic audio components.
18 . The audio system of claim 16 , wherein the first set of Ambisonic audio components includes an omnidirectional pattern (W) audio component, a first bi-directional polar pattern audio component aligned in a first direction (X), a second bi-directional polar pattern audio component aligned in a second direction (Y), and a third bi-directional polar pattern audio component aligned in a third direction (Z), and the first layer of the bit stream does not contain other encoded Ambisonic audio components.
19 . A processor to execute instructions stored in a non-transitory computer-readable memory, causing the processor to:
generate a first set of Ambisonic audio components based on ambience and one or more object-based audio signals; encode into a bit stream, a first layer having the first set of Ambisonic audio components; and encode into the bit stream, a second layer having at least one of the one or more object-based audio signals.
20 . The processor of claim 19 , wherein the first set of Ambisonic audio components includes an omnidirectional pattern (W) audio component, first order Ambisonic audio components, and second order Ambisonic audio components, and the first layer of the bit stream does not contain other encoded Ambisonic audio components.Join the waitlist — get patent alerts
Track US2022262373A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.