Sound modification using machine learing classification
Abstract
A method includes receiving input from a user that identifies types of audio as positive audio or negative audio. The method further includes outputting, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio. The method further includes identifying, with a microphone, audio in a physical environment. The method further includes splitting, with an audio machine-learning model, the audio into audio sources. The method further includes outputting, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio. The method further includes modifying output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A computer-implemented method comprising:
receiving input from a user that identifies types of audio as positive audio or negative audio; outputting, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio; identifying, with a microphone, audio in a physical environment; splitting, with an audio machine-learning model, the audio into audio sources; outputting, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and modifying output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.
2 . The method of claim 1 , further comprising:
responsive to receiving the input from the user, generating a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and generating a user profile based on the input from the user.
3 . The method of claim 2 , further comprising:
detecting, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound; classifying, with the large language model, a particular type of audio based on the instruction from the user; splitting, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and modifying the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.
4 . The method of claim 3 , further comprising:
updating the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.
5 . The method of claim 1 , further comprising:
detecting, with the microphone, an instruction from a user to increase a volume of a person; classifying, with the large language model, a particular type of audio based on the instruction from the user; and modifying output of the auditory device to increase the volume of the person.
6 . The method of claim 5 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person.
7 . The method of claim 1 , wherein the large language model is trained by:
receiving training data that includes audio files and descriptions of the audio files; outputting, with an audio encoder, embedded audio; generating audio tokens; tokenizing the descriptions of the audio files; outputting, with a text embedding layer, text tokens; and providing pairs of audio tokens and corresponding text tokens to the large language model for training.
8 . The method of claim 1 , wherein the positive audio source is amplified and the negative audio source is reduced or cancelled based on a degree of amplification or reduction input by the user.
9 . A system comprising:
one or more processors; and logic encoded in one or more non-transitory media for execution by the one or more processors and when executed are operable to:
receive input from a user that identifies types of audio as positive audio or negative audio;
output, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio;
identify, with a microphone, audio in a physical environment;
split, with an audio machine-learning model, the audio into audio sources;
output, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and
modify output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.
10 . The system of claim 9 , wherein the logic is further operable to:
responsive to receiving the input from the user, generate a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and generate a user profile based on the input from the user.
11 . The system of claim 10 , wherein the logic is further operable to:
detect, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound; classify, with the large language model, a particular type of audio based on the instruction from the user; split, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and modify the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.
12 . The system of claim 11 , wherein the logic is further operable to:
update the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.
13 . The system of claim 9 , wherein the logic is further operable to:
detect, with the microphone, an instruction from a user to increase a volume of a person; classify, with the large language model, a particular type of audio based on the instruction from the user; and modify output of the auditory device to increase the volume of the person.
14 . The system of claim 13 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person.
15 . Software encoded in one or more computer-readable media for execution by one or more processors of an auditory device and when executed is operable to:
receive input from a user that identifies types of audio as positive audio or negative audio; output, with a large language model, a classification of the types of audio that the user identified as positive audio or negative audio; identify, with a microphone, audio in a physical environment; split, with an audio machine-learning model, the audio into audio sources; output, with the large language model, one or more matched audio sources that are matched with the one or more types of audio that the user categorized as positive audio or negative audio; and modify output of an auditory device based on the one or more matched audio sources, wherein a positive audio source is amplified and a negative audio source is reduced or cancelled.
16 . The software of claim 15 , wherein the software is further operable to:
responsive to receiving the input from the user, generate a user interface that includes a list of the types of audio that the user identified as positive audio or negative audio, wherein the user interface includes options for specifying a degree of positive audio or a degree of negative audio; and generate a user profile based on the input from the user.
17 . The software of claim 16 , wherein the software is further operable to:
detect, with the microphone, an instruction from a user to increase or decrease a volume of a type of sound; classify, with the large language model, a particular type of audio based on the instruction from the user; split, with the audio machine-learning model, the audio in the physical environment to isolate the particular type of audio; and modify the output of the auditory device based on the instruction from the user to increase or decrease the volume of particular type of audio.
18 . The software of claim 17 , wherein the software is further operable to:
update the user profile to categorize the type of sound as positive audio or negative audio based on the instruction from the user.
19 . The software of claim 15 , wherein the software is further operable to:
detect, with the microphone, an instruction from a user to increase a volume of a person; classify, with the large language model, a particular type of audio based on the instruction from the user; and modify output of the auditory device to increase the volume of the person.
20 . The software of claim 19 , wherein the instruction from the user further includes an identification of a location of the person as identified by the large language model and classifying the particular type of audio is further based on the location of the person.Join the waitlist — get patent alerts
Track US2025279111A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.