US11727922B2ActiveUtilityA1

Systems and methods for deriving expression of intent from recorded speech

Assignee: VERINT AMERICAS INCPriority: Jan 17, 2017Filed: May 11, 2021Granted: Aug 15, 2023
Est. expiryJan 17, 2037(~10.5 yrs left)· nominal 20-yr term from priority
Inventors:Moshe Villaizan
G10L 15/1815G10L 15/02G10L 25/51G10L 2015/088G10L 15/26G10L 15/1822G10L 15/187
87
PatentIndex Score
4
Cited by
19
References
20
Claims

Abstract

A computerized system for deriving expression of intent from recorded speech includes: a text classification module comparing a transcription of recorded speech against a text classifier to generate a first set of representations of potential intents; a phonetics classification module comparing a phonetic transcription of the recorded speech against a phonetics classifier to generate a second set of representations; an audio classification module comparing an audio version of the recorded speech with an audio classifier to generate a third set of representations; and a discriminator module for receiving the first, second and third sets of the representations of potential intents and generating one derived expression of intent by processing the first, second and third sets together; where at least two of the text classification module, the phonetics classification module, and the audio classification module are asynchronous processes from one another.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A computerized system for transforming recorded speech into a derived expression of intent from the recorded speech, the computerized system comprising:
 a speech-to-text normalizer computer comprising a plurality of programming instructions stored in a memory and operating on a processor, the plurality of programming instructions when executed by the processor cause the processor to:
 receive an audio file; 
 convert the audio file to a converted single-channel audio file format; 
 automatically transcribe the converted single-channel audio file format to produce a speech transcript comprising a written transcription of a selection of recorded speech within the single-channel formatted audio file; 
 automatically analyze the speech transcript using a plurality of text classifiers to produce a plurality of labeled text classification results; 
 assign a label and a confidence value, of a plurality of confidence values, to at least a portion of the speech transcript, wherein the confidence value is assigned based upon a comparison of the portion of the speech transcript, developed from one or more phonemes, with key words established in one or more of a dictionary knowledge base, a vocabulary knowledge base, and a language model knowledge base, and wherein the plurality of labeled text classification results are produced based on a comparison of the confidence value assigned by a text classifier, of the plurality of text classifiers, against a stored confidence threshold; 
 process an audio classification, using an audio classifier, for at least a portion of the audio file using a time-domain pitch synchronous overlap to produce an audio output sample comprising an audio segment wherein the pitch and duration of the portion of the audio file have been modified, wherein the audio classifier is trained by:
 using a speech synthesizer module to generate a plurality of machine-synthesized audio speech signals that are each associated with a known intent for a given signal, 
 extracting one or more features from the machine-synthesized audio speech signals to generate a feature pattern for each machine-synthesized audio speech signal, 
 generating an audio signal classifier training set based on the feature patterns, and 
 training the audio classifier using the audio signal classifier training set; 
 
 analyze a combined input to determine an expression of intent based on a comparison of the labels and the plurality of confidence values within the plurality of labeled text classification results and the pitch and duration of an audio segment within the audio output sample, the combined input comprising the plurality of labeled text classification results and the portion of the audio file, 
 wherein at least one of the plurality of labeled text classification results and the audio classification are asynchronous processes from one another; and 
 output the determined expression of intent. 
 
 
     
     
       2. The computerized system of  claim 1 , further comprising one or more databases comprising:
 a plurality of potential intents that are derived from the portion of the audio file, wherein for each potential intent, of the plurality of potential intents, the key words correspond to a plurality of different potential expressions; and 
 a plurality of representations associated with the plurality of potential intents derived from the portion of the audio file, wherein for each potential intent, audio features for machine generated audio signals previously generated for a corresponding plurality of different potential expressions for each potential intent. 
 
     
     
       3. The computerized system of  claim 2  wherein the plurality of programming instructions when further executed by the processor cause the processor to:
 calculate one or more confidence scores associated with at least one set of the plurality of potential intents; and 
 generate at least one derived expression of intent by processing at least one set of the plurality of potential intents and associated one or more confidence scores together. 
 
     
     
       4. The computerized system of  claim 1 , wherein at least one of the plurality of text classifiers and the audio classifier lie on a parallel processing path with one another. 
     
     
       5. The computerized system of  claim 1 , wherein the plurality of programming instructions when further executed by the processor cause the processor to generate a keyword-based transcription from the portion of the audio file for text classification. 
     
     
       6. The computerized system of  claim 1 , wherein the plurality of programming instructions when further executed by the processor cause the processor to:
 analyze the speech transcript using a plurality of phonetic classifiers to produce a phonetic transcript, comprising the pronunciation of each of the words contained within the speech transcript, wherein each of the plurality of phonetic classifiers indexes a plurality of words according to their pronunciation; and 
 provide the phonetic transcript as output. 
 
     
     
       7. The computerized system of  claim 6 , further comprising one or more databases comprising:
 the one or more phenomes, previously derived by phonetics classification, using a phonetics classifier from the plurality of phonetic classifiers, for the key words correspond to a plurality of different potential expressions for each potential intent. 
 
     
     
       8. The computerized system of  claim 6 , wherein the phonetics classification and the audio classification are asynchronous processes from one another. 
     
     
       9. The computerized system of  claim 8 , wherein the plurality of programming instructions when further executed by the processor cause the processor to process the text classification, the phonetics classification, and the audio classification on parallel processing paths. 
     
     
       10. The computerized system of  claim 6 , wherein the plurality of programming instructions when further executed by the processor cause the processor to generate a keyword-based transcription from the portion of the audio file. 
     
     
       11. The computerized system of  claim 6 , wherein the plurality of programming instructions when further executed by the processor cause the processor to:
 calculate confidence scores associated with the at least one set of the plurality of potential intents; and 
 generate at least one derived expression of intent by processing the at least one set of the plurality of potential intents and associated confidence scores together. 
 
     
     
       12. A computer-implemented method for transforming recorded speech into a derived expression of intent from the recorded speech, the method comprising:
 receiving, by a speech-to-text normalizer computer, an audio file; 
 converting, by the speech-to-text normalizer computer, the audio file to a converted single-channel audio file format; 
 automatically transcribing, by the speech-to-text normalizer computer, the converted single-channel audio file format to produce a speech transcript comprising a written transcription of a selection of recorded speech within the single-channel formatted audio file; 
 automatically analyzing, by the speech-to-text normalizer computer, the speech transcript using a plurality of text classifiers to produce a plurality of labeled text classification results; 
 assigning, by the speech-to-text normalizer computer, a label and a confidence value, of a plurality of confidence values, to at least a portion of the speech transcript, wherein the confidence value is assigned based upon a comparison of the portion of the speech transcript developed from one or more phonemes with key words established in one or more of a dictionary knowledge base, a vocabulary knowledge base, and a language model knowledge base, and wherein the plurality of labeled text classification results are produced based on a comparison of the confidence value assigned by a text classifier against a stored confidence threshold;
 processing, by the speech-to-text normalizer computer and using an audio classifier, an audio classification wherein the at least a portion of the audio file using a time-domain pitch synchronous overlap to produce an audio output sample comprising an audio segment wherein the pitch and duration of the portion of the audio file have been modified wherein the audio classifier is trained by:
 using a speech synthesizer module to generate a plurality of machine-synthesized audio speech signals that are each associated with a known intent for a given signal, 
 extracting one or more features from the machine-synthesized audio speech signals to generate a feature pattern for each machine-synthesized audio speech signal, 
 generating an audio signal classifier training set based on the feature patterns, and 
 training the audio classifier using the audio signal classifier training set; 
 
 
 and 
 analyzing, by the speech-to-text normalizer computer, a combined input to determine an expression of intent based on a comparison of the labels and the plurality of confidence values within the plurality of labeled text classification results and the pitch and duration of an audio segment within the audio output sample, the combined input comprising the plurality of labeled text classification results and the portion of the audio file, 
 wherein at least one of the plurality of labeled text classification results and the audio classification are asynchronous processes from one another; and 
 outputting the determined expression of intent. 
 
     
     
       13. The method of  claim 12 , further comprising:
 storing, by the speech-to-text normalizer computer, in one or more databases: 
 a plurality of potential intents that are derived from the portion of the audio file, wherein for each potential intent, of the plurality of potential intents, the key words correspond to a plurality of different potential expressions; and 
 a plurality of representations associated with the plurality of potential intents derived from the portion of the audio file, wherein for each potential intent, audio features for machine generated audio signals previously generated for a corresponding plurality of different potential expressions for each potential intent. 
 
     
     
       14. The method of  claim 13 , further comprising:
 calculating, by the speech-to-text normalizer computer, one or more confidence scores associated with at least one set of the plurality of potential intents; and 
 generating, by the speech-to-text normalizer computer, at least one derived expression of intent by processing at least one set of the plurality of potential intents and associated one or more confidence scores together. 
 
     
     
       15. The method of  claim 12 , wherein at least one of the text classification and the audio classification are processed, by the speech-to-text normalizer computer, on a parallel processing path with one another. 
     
     
       16. The method of  claim 12 , further comprising generating, by the speech-to-text normalizer computer, a keyword-based transcription from the at least portion of the audio file for text classification. 
     
     
       17. The method of  claim 12 , further comprising:
 analyzing, by the speech-to-text normalizer computer, the speech transcript using a plurality of phonetic classifiers to produce a phonetic transcript, comprising the pronunciation of each of the words contained within the speech transcript, wherein each of the plurality of phonetic classifiers indexes a plurality of words according to their pronunciation; and 
 providing, by the speech-to-text normalizer computer, the phonetic transcript as output. 
 
     
     
       18. The method of  claim 17 , wherein the text classification, the phonetics classification, and the audio classification are performed, by the speech-to-text normalizer computer, on parallel processing paths. 
     
     
       19. The method of  claim 17 , further comprising generating, by the speech-to-text normalizer computer, a keyword-based transcription from the at least a portion of the audio file. 
     
     
       20. The method of  claim 17 , further comprising:
 calculating, by the speech-to-text normalizer computer, confidence scores associated with the at least one set of the plurality of potential intents; and 
 generating, by the speech-to-text normalizer computer, at least one derived expression of intent by processing the at least one set of the plurality of potential intents and associated confidence scores together.

Join the waitlist — get patent alerts

Track US11727922B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.