Audio processing for voice simulated noise effects
Abstract
Systems and methods may be used to process and output information related to a non-speech vocalization, for example from a user attempting to mimic a non-speech sound. A method may include determine a mimic quality value associated with an audio file by comparing a non-speech vocalization to a prerecorded audio file. For example, the method may include determining an edit distance between the non-speech vocalization and the prerecorded audio file. The method may include assigning a mimic quality value to the audio file based on the edit distance. The method may include outputting the mimic quality value.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A device comprising:
a display to provide a user interface for interacting with a social bot; memory; and a processor in communication with the memory, the processor to:
provide an indication initiating an impression game within the user interface with the social bot, the indication indicating a non-speech sound to be mimicked;
receive an audio file or streamed audio including a non-speech vocalization from a user attempting to mimic the non-speech sound via the user interface;
determine a mimic quality value associated with the audio file or the streamed audio by comparing the non-speech vocalization to a prerecorded audio file in a database; and
output a response to the received audio file or the streamed audio from the social bot for display on the user interface based on the mimic quality value.
2 . The device of claim 1 , wherein the prerecorded audio file is a recording of the non-speech sound to be mimicked or a recording of a person mimicking the non-speech sound.
3 . The device of claim 1 , wherein the processor is further to provide a token via the user interface in response to the mimic quality value exceeding a threshold, the token used to unlock digital content.
4 . The device of claim 1 , wherein to determine the mimic quality value, the processor is further to determine whether the non-speech vocalization is within a predetermined edit distance of the prerecorded audio file.
5 . The device of claim 4 , wherein the response is positive when the non-speech vocalization is within the predetermined edit distance.
6 . The device of claim 4 , wherein to determine whether the non-speech vocalization is within the predetermined edit distance of the prerecorded audio file, the processor is further to use dynamic time warping.
7 . The device of claim 4 , wherein to determine whether the non-speech vocalization is within the predetermined edit distance of the prerecorded audio file, the processor is further to use Mel Frequency Cepstrum Coefficients representing the audio file or the streamed audio and the prerecorded audio file to compare frames of the audio file or the streamed audio to frames of the prerecorded audio file to determine an edit distance between the audio file or the streamed audio and the prerecorded audio file.
8 . The device of claim 7 , wherein the Mel Frequency Cepstrum Coefficients are generated by performing a fast fourier transform on the audio file or the streamed audio and the prerecorded audio file, mapping results of the fast fourier transform to a Mel scale, and determining amplitudes of the results mapped to the Mel scale, including a first series of amplitudes corresponding to the audio file or the streamed audio and a second series of amplitudes corresponding to the prerecorded audio file.
9 . The device of claim 4 , wherein the edit distance is a number of changes, edits, or deletions needed to convert the audio file or the streamed audio to the prerecorded audio file.
10 . A method comprising:
using a processor in communication with memory to:
provide an indication initiating an impression game within a user interface on a display with the social bot, the indication indicating a non-speech sound to be mimicked;
receive an audio file or streamed audio including a non-speech vocalization from a user attempting to mimic the non-speech sound via the user interface;
determine a mimic quality value associated with the audio file or the streamed audio by comparing the non-speech vocalization to a prerecorded audio file in a database; and
output a response to the received audio file or the streamed audio from the social bot for display on the user interface based on the mimic quality value.
11 . The method of claim 10 , wherein to determine the mimic quality value, the processor is to compare the non-speech vocalization to a plurality of prerecorded audio files in the database.
12 . The method of claim 10 , wherein the prerecorded audio file is a recording of the non-speech sound to be mimicked or of a person mimicking the non-speech sound.
13 . The method of claim 10 , further comprising using the processor to generate an auditory interaction mimicking a second non-speech sound to be presented from the social bot via the user interface.
14 . The method of claim 13 , further comprising using the processor to receive a user guess of the second non-speech sound in the auditory interaction, and provide, from the social bot, a response to the user guess via the user interface, wherein the user guess includes at least one of a text response, a spoken response, an emoji response, an emoticon response, or an image response.
15 . The method of claim 10 , wherein to determine the mimic quality value associated with the audio file includes using a machine learning classifier to determine whether the non-speech vocalization matches the prerecorded audio file.
16 . The method of claim 13 , further comprising using the processor to present the auditory interaction via the user interface from the social bot and a contextual clue related to the second non-speech sound.
17 . The method of claim 10 , further comprising using the processor to detect a spoken word in the audio file or the streamed audio and use the spoken word to determine the mimic quality value.
18 . The method of claim 10 , wherein to compare the non-speech vocalization to the prerecorded audio file includes using the processor to compare an extracted speech portion of the audio file or the streamed audio to a speech portion of the prerecorded audio file.
19 . At least one non-transitory machine-readable medium including instructions for performing operations, which when executed by a processor, cause the processor to:
provide an indication initiating an impression game within a user interface on a display with the social bot, the indication indicating a non-speech sound to be mimicked, receive an audio file or streamed audio including a non-speech vocalization from a user attempting to mimic the non-speech sound via the user interface; determine a mimic quality value associated with the audio file or the streamed audio by comparing the non-speech vocalization to a prerecorded audio file in a database; and output a response to the received audio file or the streamed audio from the social bot for display on the user interface based on the mimic quality value.
20 . The at least one machine-readable medium of claim 19 , wherein the database is a structured database of prerecorded audio files arranged by non-speech sound, and wherein to compare the non-speech vocalization to the prerecorded audio file, the instructions further cause the processor to select the prerecorded audio file from the structured database based on the non-speech sound to be mimicked indicated in the indication.Join the waitlist — get patent alerts
Track US2019109804A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.