US2025279111A1PendingUtilityA1

Sound modification using machine learing classification

Assignee: SONY GROUP CORPPriority: Mar 4, 2024Filed: Mar 4, 2024Published: Sep 4, 2025
Est. expiryMar 4, 2044(~17.6 yrs left)· nominal 20-yr term from priority
H04R 1/1083H04R 2430/01H04R 25/507H04R 2225/41G10L 2015/223G10L 25/51G10L 21/028G10L 15/22G10L 15/183G10L 15/063G06F 3/165G10L 21/0316G10L 21/0272G06F 40/30G06F 3/167G10L 21/034
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes receiving input from a user that identifies types of audio as positive audio or negative audio. The method further includes outputting, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio. The method further includes identifying, with a microphone, audio in a physical environment. The method further includes splitting, with an audio machine-learning model, the audio into audio sources. The method further includes outputting, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio. The method further includes modifying output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A computer-implemented method comprising:
 receiving input from a user that identifies types of audio as positive audio or negative audio;   outputting, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio;   identifying, with a microphone, audio in a physical environment;   splitting, with an audio machine-learning model, the audio into audio sources;   outputting, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and   modifying output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.   
     
     
         2 . The method of  claim 1 , further comprising:
 responsive to receiving the input from the user, generating a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and   generating a user profile based on the input from the user.   
     
     
         3 . The method of  claim 2 , further comprising:
 detecting, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound;   classifying, with the large language model, a particular type of audio based on the instruction from the user;   splitting, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and   modifying the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.   
     
     
         4 . The method of  claim 3 , further comprising:
 updating the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.   
     
     
         5 . The method of  claim 1 , further comprising:
 detecting, with the microphone, an instruction from a user to increase a volume of a person;   classifying, with the large language model, a particular type of audio based on the instruction from the user; and   modifying output of the auditory device to increase the volume of the person.   
     
     
         6 . The method of  claim 5 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person. 
     
     
         7 . The method of  claim 1 , wherein the large language model is trained by:
 receiving training data that includes audio files and descriptions of the audio files;   outputting, with an audio encoder, embedded audio;   generating audio tokens;   tokenizing the descriptions of the audio files;   outputting, with a text embedding layer, text tokens; and   providing pairs of audio tokens and corresponding text tokens to the large language model for training.   
     
     
         8 . The method of  claim 1 , wherein the positive audio source is amplified and the negative audio source is reduced or cancelled based on a degree of amplification or reduction input by the user. 
     
     
         9 . A system comprising:
 one or more processors; and   logic encoded in one or more non-transitory media for execution by the one or more processors and when executed are operable to:
 receive input from a user that identifies types of audio as positive audio or negative audio; 
 output, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio; 
 identify, with a microphone, audio in a physical environment; 
 split, with an audio machine-learning model, the audio into audio sources; 
 output, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and 
 modify output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled. 
   
     
     
         10 . The system of  claim 9 , wherein the logic is further operable to:
 responsive to receiving the input from the user, generate a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and   generate a user profile based on the input from the user.   
     
     
         11 . The system of  claim 10 , wherein the logic is further operable to:
 detect, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound;   classify, with the large language model, a particular type of audio based on the instruction from the user;   split, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and   modify the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.   
     
     
         12 . The system of  claim 11 , wherein the logic is further operable to:
 update the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.   
     
     
         13 . The system of  claim 9 , wherein the logic is further operable to:
 detect, with the microphone, an instruction from a user to increase a volume of a person;   classify, with the large language model, a particular type of audio based on the instruction from the user; and   modify output of the auditory device to increase the volume of the person.   
     
     
         14 . The system of  claim 13 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person. 
     
     
         15 . Software encoded in one or more computer-readable media for execution by one or more processors of an auditory device and when executed is operable to:
 receive input from a user that identifies types of audio as positive audio or negative audio;   output, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio;   identify, with a microphone, audio in a physical environment;   split, with an audio machine-learning model, the audio into audio sources;   output, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and   modify output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.   
     
     
         16 . The software of  claim 15 , wherein the software is further operable to:
 responsive to receiving the input from the user, generate a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and   generate a user profile based on the input from the user.   
     
     
         17 . The software of  claim 16 , wherein the software is further operable to:
 detect, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound;   classify, with the large language model, a particular type of audio based on the instruction from the user;   split, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and   modify the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.   
     
     
         18 . The software of  claim 17 , wherein the software is further operable to:
 update the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.   
     
     
         19 . The software of  claim 15 , wherein the software is further operable to:
 detect, with the microphone, an instruction from a user to increase a volume of a person;   classify, with the large language model, a particular type of audio based on the instruction from the user; and   modify output of the auditory device to increase the volume of the person.   
     
     
         20 . The software of  claim 19 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person.

Join the waitlist — get patent alerts

Track US2025279111A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.