Systems and methods for deriving expression of intent from recorded speech
Abstract
A computerized system for deriving expression of intent from recorded speech includes: a text classification module comparing a transcription of recorded speech against a text classifier to generate a first set of representations of potential intents; a phonetics classification module comparing a phonetic transcription of the recorded speech against a phonetics classifier to generate a second set of representations; an audio classification module comparing an audio version of the recorded speech with an audio classifier to generate a third set of representations; and a discriminator module for receiving the first, second and third sets of the representations of potential intents and generating one derived expression of intent by processing the first, second and third sets together; where at least two of the text classification module, the phonetics classification module, and the audio classification module are asynchronous processes from one another.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A computerized system for transforming recorded speech into a derived expression of intent from the recorded speech, the computerized system comprising:
a speech-to-text normalizer computer comprising a plurality of programming instructions stored in a memory and operating on a processor, the plurality of programming instructions when executed by the processor cause the processor to:
receive an audio file;
convert the audio file to a converted single-channel audio file format;
automatically transcribe the converted single-channel audio file format to produce a speech transcript comprising a written transcription of a selection of recorded speech within the single-channel formatted audio file;
automatically analyze the speech transcript using a plurality of text classifiers to produce a plurality of labeled text classification results;
assign a label and a confidence value, of a plurality of confidence values, to at least a portion of the speech transcript, wherein the confidence value is assigned based upon a comparison of the portion of the speech transcript, developed from one or more phonemes, with key words established in one or more of a dictionary knowledge base, a vocabulary knowledge base, and a language model knowledge base, and wherein the plurality of labeled text classification results are produced based on a comparison of the confidence value assigned by a text classifier, of the plurality of text classifiers, against a stored confidence threshold;
process an audio classification, using an audio classifier, for at least a portion of the audio file using a time-domain pitch synchronous overlap to produce an audio output sample comprising an audio segment wherein the pitch and duration of the portion of the audio file have been modified, wherein the audio classifier is trained by:
using a speech synthesizer module to generate a plurality of machine-synthesized audio speech signals that are each associated with a known intent for a given signal,
extracting one or more features from the machine-synthesized audio speech signals to generate a feature pattern for each machine-synthesized audio speech signal,
generating an audio signal classifier training set based on the feature patterns, and
training the audio classifier using the audio signal classifier training set;
analyze a combined input to determine an expression of intent based on a comparison of the labels and the plurality of confidence values within the plurality of labeled text classification results and the pitch and duration of an audio segment within the audio output sample, the combined input comprising the plurality of labeled text classification results and the portion of the audio file,
wherein at least one of the plurality of labeled text classification results and the audio classification are asynchronous processes from one another; and
output the determined expression of intent.
2. The computerized system of claim 1 , further comprising one or more databases comprising:
a plurality of potential intents that are derived from the portion of the audio file, wherein for each potential intent, of the plurality of potential intents, the key words correspond to a plurality of different potential expressions; and
a plurality of representations associated with the plurality of potential intents derived from the portion of the audio file, wherein for each potential intent, audio features for machine generated audio signals previously generated for a corresponding plurality of different potential expressions for each potential intent.
3. The computerized system of claim 2 wherein the plurality of programming instructions when further executed by the processor cause the processor to:
calculate one or more confidence scores associated with at least one set of the plurality of potential intents; and
generate at least one derived expression of intent by processing at least one set of the plurality of potential intents and associated one or more confidence scores together.
4. The computerized system of claim 1 , wherein at least one of the plurality of text classifiers and the audio classifier lie on a parallel processing path with one another.
5. The computerized system of claim 1 , wherein the plurality of programming instructions when further executed by the processor cause the processor to generate a keyword-based transcription from the portion of the audio file for text classification.
6. The computerized system of claim 1 , wherein the plurality of programming instructions when further executed by the processor cause the processor to:
analyze the speech transcript using a plurality of phonetic classifiers to produce a phonetic transcript, comprising the pronunciation of each of the words contained within the speech transcript, wherein each of the plurality of phonetic classifiers indexes a plurality of words according to their pronunciation; and
provide the phonetic transcript as output.
7. The computerized system of claim 6 , further comprising one or more databases comprising:
the one or more phenomes, previously derived by phonetics classification, using a phonetics classifier from the plurality of phonetic classifiers, for the key words correspond to a plurality of different potential expressions for each potential intent.
8. The computerized system of claim 6 , wherein the phonetics classification and the audio classification are asynchronous processes from one another.
9. The computerized system of claim 8 , wherein the plurality of programming instructions when further executed by the processor cause the processor to process the text classification, the phonetics classification, and the audio classification on parallel processing paths.
10. The computerized system of claim 6 , wherein the plurality of programming instructions when further executed by the processor cause the processor to generate a keyword-based transcription from the portion of the audio file.
11. The computerized system of claim 6 , wherein the plurality of programming instructions when further executed by the processor cause the processor to:
calculate confidence scores associated with the at least one set of the plurality of potential intents; and
generate at least one derived expression of intent by processing the at least one set of the plurality of potential intents and associated confidence scores together.
12. A computer-implemented method for transforming recorded speech into a derived expression of intent from the recorded speech, the method comprising:
receiving, by a speech-to-text normalizer computer, an audio file;
converting, by the speech-to-text normalizer computer, the audio file to a converted single-channel audio file format;
automatically transcribing, by the speech-to-text normalizer computer, the converted single-channel audio file format to produce a speech transcript comprising a written transcription of a selection of recorded speech within the single-channel formatted audio file;
automatically analyzing, by the speech-to-text normalizer computer, the speech transcript using a plurality of text classifiers to produce a plurality of labeled text classification results;
assigning, by the speech-to-text normalizer computer, a label and a confidence value, of a plurality of confidence values, to at least a portion of the speech transcript, wherein the confidence value is assigned based upon a comparison of the portion of the speech transcript developed from one or more phonemes with key words established in one or more of a dictionary knowledge base, a vocabulary knowledge base, and a language model knowledge base, and wherein the plurality of labeled text classification results are produced based on a comparison of the confidence value assigned by a text classifier against a stored confidence threshold;
processing, by the speech-to-text normalizer computer and using an audio classifier, an audio classification wherein the at least a portion of the audio file using a time-domain pitch synchronous overlap to produce an audio output sample comprising an audio segment wherein the pitch and duration of the portion of the audio file have been modified wherein the audio classifier is trained by:
using a speech synthesizer module to generate a plurality of machine-synthesized audio speech signals that are each associated with a known intent for a given signal,
extracting one or more features from the machine-synthesized audio speech signals to generate a feature pattern for each machine-synthesized audio speech signal,
generating an audio signal classifier training set based on the feature patterns, and
training the audio classifier using the audio signal classifier training set;
and
analyzing, by the speech-to-text normalizer computer, a combined input to determine an expression of intent based on a comparison of the labels and the plurality of confidence values within the plurality of labeled text classification results and the pitch and duration of an audio segment within the audio output sample, the combined input comprising the plurality of labeled text classification results and the portion of the audio file,
wherein at least one of the plurality of labeled text classification results and the audio classification are asynchronous processes from one another; and
outputting the determined expression of intent.
13. The method of claim 12 , further comprising:
storing, by the speech-to-text normalizer computer, in one or more databases:
a plurality of potential intents that are derived from the portion of the audio file, wherein for each potential intent, of the plurality of potential intents, the key words correspond to a plurality of different potential expressions; and
a plurality of representations associated with the plurality of potential intents derived from the portion of the audio file, wherein for each potential intent, audio features for machine generated audio signals previously generated for a corresponding plurality of different potential expressions for each potential intent.
14. The method of claim 13 , further comprising:
calculating, by the speech-to-text normalizer computer, one or more confidence scores associated with at least one set of the plurality of potential intents; and
generating, by the speech-to-text normalizer computer, at least one derived expression of intent by processing at least one set of the plurality of potential intents and associated one or more confidence scores together.
15. The method of claim 12 , wherein at least one of the text classification and the audio classification are processed, by the speech-to-text normalizer computer, on a parallel processing path with one another.
16. The method of claim 12 , further comprising generating, by the speech-to-text normalizer computer, a keyword-based transcription from the at least portion of the audio file for text classification.
17. The method of claim 12 , further comprising:
analyzing, by the speech-to-text normalizer computer, the speech transcript using a plurality of phonetic classifiers to produce a phonetic transcript, comprising the pronunciation of each of the words contained within the speech transcript, wherein each of the plurality of phonetic classifiers indexes a plurality of words according to their pronunciation; and
providing, by the speech-to-text normalizer computer, the phonetic transcript as output.
18. The method of claim 17 , wherein the text classification, the phonetics classification, and the audio classification are performed, by the speech-to-text normalizer computer, on parallel processing paths.
19. The method of claim 17 , further comprising generating, by the speech-to-text normalizer computer, a keyword-based transcription from the at least a portion of the audio file.
20. The method of claim 17 , further comprising:
calculating, by the speech-to-text normalizer computer, confidence scores associated with the at least one set of the plurality of potential intents; and
generating, by the speech-to-text normalizer computer, at least one derived expression of intent by processing the at least one set of the plurality of potential intents and associated confidence scores together.Join the waitlist — get patent alerts
Track US11727922B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.