US2023223035A1PendingUtilityA1
Systems and methods for visually guided audio separation
Est. expiryDec 6, 2039(~13.4 yrs left)· nominal 20-yr term from priority
G06N 3/0895G06N 3/0464G10L 21/028G10L 25/30G06T 7/11G06T 7/168G06T 7/174G06T 2207/20081G06T 2207/20084G06T 2207/10016G06T 2207/20056G06N 3/063G06N 3/084G06N 3/045G06T 7/70G10L 25/51G06N 3/04G10L 25/18G10L 25/57
63
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system for separating audio based on sound producing objects includes a processor configured to receive video data and audio data. The processor is also configured to perform object detection using the video data to identify a number of sound producing objects in the video data and predict a separation for each sound producing object detected in the video data. The processor is also configured to generate separated audio data for each sound producing object using the separation and the audio data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system comprising:
a processor configured to:
detect a number of sound producing objects in video data;
determine a separation for one or more of the sound producing objects detected in the video data that increases a reliability of a separation of sounds produced by a corresponding class of sound producing objects in the video data; and
generate separated audio data for one or more of the sound producing objects using the separation and audio data associated with the video data.
2 . The system of claim 1 , wherein the processor is further configured to determine the separation to minimize a co-separation loss.
3 . The system of claim 2 , wherein the processor is configured to:
convert the audio data associated with the video data into a magnitude spectrogram; determine a spectrogram mask for each of the one or more of the sound producing objects as the separation; generate a separated spectrogram for each detected object using the spectrogram mask and the magnitude spectrogram; and convert each of the separated spectrograms to audio data to generate the separated audio data.
4 . The system of claim 3 , wherein the co-separation loss is a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask.
5 . The system of claim 4 , wherein the ground-truth spectrogram ratio mask is a ratio between:
(1) a particular magnitude spectrogram of the audio data of one of one or more sets of audio data and associated video data; and (2) a sum of magnitude spectrograms of the audio data of the one or more sets.
6 . The system of claim 1 , wherein the reliability is quantitatively related to a consistency loss, wherein the separation is determined to minimize the consistency loss to increase the reliability.
7 . The system of claim 1 , wherein the processor is further configured to use a neural network to determine the separation and generate the separated audio data for the one or more of the so and producing objects, wherein the neural network is trained using a plurality of sets of the video data and the audio data.
8 . A method comprising:
detecting a number of sound producing objects in video data; determining a separation for one or more of the sound producing objects detected in the video data that increases a reliability of a separation of sounds produced by a corresponding class of sound producing objects in the video data; and generating separated audio data for one or more of the sound producing objects using the separation and audio data associated with the video data.
9 . The method of claim 8 , wherein the separation is determined to minimize a co-separation loss.
10 . The method of claim 9 , comprising:
converting the audio data associated with the video data into a magnitude spectrogram; determining a spectrogram mask for each of the one or more of the sound producing objects as the separation; generating a separated spectrogram for each detected object using the spectrogram mask and the magnitude spectrogram; and converting each of the separated spectrograms to audio data to generate the separated audio data.
11 . The method of claim 10 , wherein the co-separation loss is a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask.
12 . The method of claim 11 , wherein the ground-truth spectrogram ratio mask is a ratio between:
(1) a particular magnitude spectrogram of the audio data of one of one or more sets of audio data and associated video data; and (2) a sum of magnitude spectrograms of the audio data of the one or more sets.
13 . The method of claim 8 , wherein the reliability is quantitatively related to a consistency loss, wherein the separation is determined to minimize the consistency loss to increase the reliability.
14 . The method of claim 8 , wherein the method includes using a neural network to determine the separation and generate the separated audio data for the one or more of the so and producing objects, wherein the neural network is trained using a plurality of sets of the video data and the audio data.
15 . A system comprising:
a processor configured to:
receive a plurality of training data, the plurality of training data comprising one or more sets of audio data and associated video data;
perform object detection on the video data of the one or more sets to detect one or more sound producing objects of the video data;
mixing the audio data of the one or more sets to generate mixed audio data; and
training a neural network using the plurality of training data to determine a separation for the one or more sound producing objects that increases a reliability of separation of sounds produced by a corresponding class of the one or more sound producing objects.
16 . The system of claim 15 , wherein the separation is a spectrogram mask and the reliability of separation is at least partially defined by a penalty associated with a difference between the spectrogram mask for an associated sound producing object and a ground-truth spectrogram ratio mask.
17 . The system of claim 16 , wherein the ground-truth spectrogram ratio mask is a ratio between:
(1) a particular magnitude spectrogram of the audio data of one of the one or more sets of audio data and associated video data; and (2) a sum of magnitude spectrograms of the audio data of the one or more sets.
18 . The system of claim 15 , wherein the reliability of separation is at least partially defined by a cross-entropy loss function that defines a consistency loss in terms of a number of sound producing objects and a plurality of classes.
19 . The system of claim 15 , wherein the reliability of separation is defined by a combined loss function that includes a co-separation loss and a consistency loss, the combined loss function comprising a weight associated with the consistency loss.
20 . The system of claim 19 , wherein the weight associated with the consistency loss is adjustable to incentivize the neural network to determine separations associated with lower consistency losses or to determine separations associated with lower co-separation losses.Join the waitlist — get patent alerts
Track US2023223035A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.