US2025279106A1PendingUtilityA1

Audio Signal Upmixer

Assignee: SAMSUNG ELECTRONICS CO LTDPriority: Mar 4, 2024Filed: Mar 4, 2025Published: Sep 4, 2025
Est. expiryMar 4, 2044(~17.6 yrs left)· nominal 20-yr term from priority
G10L 19/008H04N 21/439
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one embodiment, a method includes accessing an audio signal encoded in an m-channel format, where m is 3 or more, and determining a time frequency representation of each of the m channels. The method further includes performing (1) simultaneous multichannel surround ambience extraction and (2) primary component extraction on the m-channel audio signal, using generalized equal-levels ambience extraction; diffusing the extracted surround ambience to an n-channel audio signal, where n is greater than m; assigning each extracted primary component of the m-channel audio signal to a primary component of a channel in the n-channel audio signal; and then obtaining an upmixed n-channel audio signal by (1) performing channel-wise addition of each primary component and ambience component of the n-channel audio signal and (2) applying an inverse time frequency operation to the channel-wise addition.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 accessing an audio signal encoded in an m-channel format, wherein m is 3 or more;   determining a time frequency representation of each of the m channels;   performing (1) simultaneous multichannel surround ambience extraction and (2) primary component extraction on the m-channel audio signal, using generalized equal-levels ambience extraction;   diffusing the extracted surround ambience to an n-channel audio signal, where n is greater than m;   assigning each extracted primary component of the m-channel audio signal to a primary component of a channel in the n-channel audio signal; and   obtaining an upmixed n-channel audio signal by (1) performing channel-wise addition of each primary component and ambience component of the n-channel audio signal and (2) applying an inverse time frequency operation to the channel-wise addition.   
     
     
         2 . The method of  claim 1 , wherein the m-channel format comprises a 5-channel surround format and the n-channel format comprises an 11-channel immersive format. 
     
     
         3 . The method of  claim 2 , further comprising generating immersive rear and side channels in the n-channel audio signal using immersive rear channels extraction. 
     
     
         4 . The method of  claim 3 , further comprising extracting, using equal-levels ambience extraction, (1) frontal ambience for an overhead front left channel and for an overhead front right channel from a left channel and a right channel in the m-channel format and (2) rear ambience for an overhead rear left channel and an overhead rear right channel from a left surround channel and a right surround channel in the m-channel format. 
     
     
         5 . The method of  claim 3 , where the method occurs in response to a request to play the audio signal. 
     
     
         6 . The method of  claim 5 , wherein the method occurs substantially in real time with the request. 
     
     
         7 . The method of  claim 1 , wherein the method is performed by a smart TV. 
     
     
         8 . The method of  claim 1 , wherein the audio signal is part of a multimedia content comprising audio and one or more images; and the method further comprises:
 identifying one or more objects in at least some of the one or more images;   determining an image position of each of the identified one or more objects;   determining, for at least some of the one or more objects, a portion of the audio signal that corresponds to that object;   determining, for each of the least some one or more objects, and based on (1) the image position of the object and (2) the portion of the audio signal that corresponds to that object, a spatial rendering for the portion of the audio signal; and   enhancing the n-channel audio signal with each determined spatial rendering.   
     
     
         9 . One or more non-transitory computer readable storage media storing instructions that are operable when executed to:
 access an audio signal encoded in an m-channel format, wherein m is 3 or more;   determine a time frequency representation of each of the m channels;   perform (1) simultaneous multichannel surround ambience extraction and (2) primary component extraction on the m-channel audio signal, using generalized equal-levels ambience extraction;   diffuse the extracted surround ambience to an n-channel audio signal, where n is greater than m;   assign each extracted primary component of the m-channel audio signal to a primary component of a channel in the n-channel audio signal; and   obtain an upmixed n-channel audio signal by (1) performing channel-wise addition of each primary component and ambience component of the n-channel audio signal and (2) applying an inverse time frequency operation to the channel-wise addition.   
     
     
         10 . The media of  claim 9 , wherein the m-channel format comprises a 5-channel surround format and the n-channel format comprises an 11-channel immersive format. 
     
     
         11 . The media of  claim 10 , wherein the instructions are further operable when executed to generate immersive rear and side channels in the n-channel audio signal using immersive rear channels extraction. 
     
     
         12 . The media of  claim 11 , wherein the instructions are further operable when executed to extract, using equal-levels ambience extraction, (1) frontal ambience for an overhead front left channel and for an overhead front right channel from a left channel and a right channel in the m-channel format and (2) rear ambience for an overhead rear left channel and an overhead rear right channel from a left surround channel and a right surround channel in the m-channel format. 
     
     
         13 . The media of  claim 11 , wherein the instructions are further operable to perform the operations in response to a request to play the audio signal. 
     
     
         14 . A system comprising:
 one or more non-transitory computer readable storage media storing instructions; and one or more processors coupled to the one or more non-transitory computer readable storage media and operable to execute the instructions to:
 access an audio signal encoded in an m-channel format, wherein m is 3 or more; 
 determine a time frequency representation of each of the m channels; 
 perform (1) simultaneous multichannel surround ambience extraction and (2) primary component extraction on the m-channel audio signal, using generalized equal-levels ambience extraction; 
 diffuse the extracted surround ambience to an n-channel audio signal, where n is greater than m; 
 assign each extracted primary component of the m-channel audio signal to a primary component of a channel in the n-channel audio signal; and 
 obtain an upmixed n-channel audio signal by (1) performing channel-wise addition of each primary component and ambience component of the n-channel audio signal and (2) applying an inverse time frequency operation to the channel-wise addition. 
   
     
     
         15 . The system of  claim 14 , wherein the m-channel format comprises a 5-channel surround format and the n-channel format comprises an 11-channel immersive format. 
     
     
         16 . The system of  claim 15 , further comprising one or more processors that are coupled to the media and are operable to execute the instructions to generate immersive rear and side channels in the n-channel audio signal using immersive rear channels extraction. 
     
     
         17 . The system of  claim 16 , further comprising one or more processors that are coupled to the media and are operable to execute the instructions to extract, using equal-levels ambience extraction, (1) frontal ambience for an overhead front left channel and for an overhead front right channel from a left channel and a right channel in the m-channel format and (2) rear ambience for an overhead rear left channel and an overhead rear right channel from a left surround channel and a right surround channel in the m-channel format. 
     
     
         18 . The system of  claim 16 , wherein the one or more processors are further operable to perform the operations in response to a request to play the audio signal. 
     
     
         19 . The system of  claim 18 , wherein the one or more processors are further operable to perform the operations substantially in real time with the request. 
     
     
         20 . The system of  claim 14 , further comprising a smart TV that contains the media and the one or more processors.

Join the waitlist — get patent alerts

Track US2025279106A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.