US2023223035A1PendingUtilityA1

Systems and methods for visually guided audio separation

Assignee: META PLATFORMS TECH LLCPriority: Dec 6, 2019Filed: Mar 14, 2023Published: Jul 13, 2023
Est. expiryDec 6, 2039(~13.4 yrs left)· nominal 20-yr term from priority
G06N 3/0895G06N 3/0464G10L 21/028G10L 25/30G06T 7/11G06T 7/168G06T 7/174G06T 2207/20081G06T 2207/20084G06T 2207/10016G06T 2207/20056G06N 3/063G06N 3/084G06N 3/045G06T 7/70G10L 25/51G06N 3/04G10L 25/18G10L 25/57
63
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system for separating audio based on sound producing objects includes a processor configured to receive video data and audio data. The processor is also configured to perform object detection using the video data to identify a number of sound producing objects in the video data and predict a separation for each sound producing object detected in the video data. The processor is also configured to generate separated audio data for each sound producing object using the separation and the audio data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising:
 a processor configured to:
 detect a number of sound producing objects in video data; 
 determine a separation for one or more of the sound producing objects detected in the video data that increases a reliability of a separation of sounds produced by a corresponding class of sound producing objects in the video data; and 
 generate separated audio data for one or more of the sound producing objects using the separation and audio data associated with the video data. 
   
     
     
         2 . The system of  claim 1 , wherein the processor is further configured to determine the separation to minimize a co-separation loss. 
     
     
         3 . The system of  claim 2 , wherein the processor is configured to:
 convert the audio data associated with the video data into a magnitude spectrogram;   determine a spectrogram mask for each of the one or more of the sound producing objects as the separation;   generate a separated spectrogram for each detected object using the spectrogram mask and the magnitude spectrogram; and   convert each of the separated spectrograms to audio data to generate the separated audio data.   
     
     
         4 . The system of  claim 3 , wherein the co-separation loss is a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask. 
     
     
         5 . The system of  claim 4 , wherein the ground-truth spectrogram ratio mask is a ratio between:
 (1) a particular magnitude spectrogram of the audio data of one of one or more sets of audio data and associated video data; and   (2) a sum of magnitude spectrograms of the audio data of the one or more sets.   
     
     
         6 . The system of  claim 1 , wherein the reliability is quantitatively related to a consistency loss, wherein the separation is determined to minimize the consistency loss to increase the reliability. 
     
     
         7 . The system of  claim 1 , wherein the processor is further configured to use a neural network to determine the separation and generate the separated audio data for the one or more of the so and producing objects, wherein the neural network is trained using a plurality of sets of the video data and the audio data. 
     
     
         8 . A method comprising:
 detecting a number of sound producing objects in video data;   determining a separation for one or more of the sound producing objects detected in the video data that increases a reliability of a separation of sounds produced by a corresponding class of sound producing objects in the video data; and   generating separated audio data for one or more of the sound producing objects using the separation and audio data associated with the video data.   
     
     
         9 . The method of  claim 8 , wherein the separation is determined to minimize a co-separation loss. 
     
     
         10 . The method of  claim 9 , comprising:
 converting the audio data associated with the video data into a magnitude spectrogram;   determining a spectrogram mask for each of the one or more of the sound producing objects as the separation;   generating a separated spectrogram for each detected object using the spectrogram mask and the magnitude spectrogram; and   converting each of the separated spectrograms to audio data to generate the separated audio data.   
     
     
         11 . The method of  claim 10 , wherein the co-separation loss is a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask. 
     
     
         12 . The method of  claim 11 , wherein the ground-truth spectrogram ratio mask is a ratio between:
 (1) a particular magnitude spectrogram of the audio data of one of one or more sets of audio data and associated video data; and   (2) a sum of magnitude spectrograms of the audio data of the one or more sets.   
     
     
         13 . The method of  claim 8 , wherein the reliability is quantitatively related to a consistency loss, wherein the separation is determined to minimize the consistency loss to increase the reliability. 
     
     
         14 . The method of  claim 8 , wherein the method includes using a neural network to determine the separation and generate the separated audio data for the one or more of the so and producing objects, wherein the neural network is trained using a plurality of sets of the video data and the audio data. 
     
     
         15 . A system comprising:
 a processor configured to:
 receive a plurality of training data, the plurality of training data comprising one or more sets of audio data and associated video data; 
 perform object detection on the video data of the one or more sets to detect one or more sound producing objects of the video data; 
 mixing the audio data of the one or more sets to generate mixed audio data; and 
 training a neural network using the plurality of training data to determine a separation for the one or more sound producing objects that increases a reliability of separation of sounds produced by a corresponding class of the one or more sound producing objects. 
   
     
     
         16 . The system of  claim 15 , wherein the separation is a spectrogram mask and the reliability of separation is at least partially defined by a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask. 
     
     
         17 . The system of  claim 16 , wherein the ground-truth spectrogram ratio mask is a ratio between:
 (1) a particular magnitude spectrogram of the audio data of one of the one or more sets of audio data and associated video data; and   (2) a sum of magnitude spectrograms of the audio data of the one or more sets.   
     
     
         18 . The system of  claim 15 , wherein the reliability of separation is at least partially defined by a cross-entropy loss function that defines a consistency loss in terms of a number of sound producing objects and a plurality of classes. 
     
     
         19 . The system of  claim 15 , wherein the reliability of separation is defined by a combined loss function that includes a co-separation loss and a consistency loss, the combined loss function comprising a weight associated with the consistency loss. 
     
     
         20 . The system of  claim 19 , wherein the weight associated with the consistency loss is adjustable to incentivize the neural network to determine separations associated with lower consistency losses or to determine separations associated with lower co-separation losses.

Join the waitlist — get patent alerts

Track US2023223035A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.