US2021335351A1PendingUtilityA1

Voice Characterization-Based Natural Language Filtering

Assignee: SOUNDHOUND INCPriority: Dec 23, 2016Filed: Jul 2, 2021Published: Oct 28, 2021
Est. expiryDec 23, 2036(~10.4 yrs left)· nominal 20-yr term from priority
Inventors:Karl Stahl
G10L 2015/025G10L 15/19G10L 15/1822G10L 17/02G10L 25/63G10L 15/1807
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An utterance is analyzed to determine a characteristic of the utterance and a transcription hypothesis is generated for the utterance. Grammar rules are then used to parse the transcription hypothesis to produce a plurality of interpretation hypotheses, each having a likelihood score. A set of authorized domains is determined based on the characteristic and the plurality of interpretation hypotheses are filtered according to the set of authorized domains. Of the remaining interpretation hypotheses, one is selected according to their likelihood scores. The characteristic may include one or more characteristics such as mood, prosody, or whether the utterance has a rising intonation.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory computer-readable medium comprising code effective to cause one or more processors to:
 characterize a speech utterance to determine at least one characteristic;   determine from the speech utterance a plurality of phoneme sequence hypotheses;   prune the plurality of phoneme sequence hypotheses to a set of phoneme sequence hypotheses conditioned on the at least one characteristic;   tokenize the set of phoneme sequence hypotheses to create word sequence hypotheses from the pruned set of phoneme sequence hypotheses;   compute probabilities for the word sequence hypotheses according to a statistical language model; and   choose as a transcription a word sequence hypothesis having the highest probability.   
     
     
         2 . The non-transitory computer-readable medium of  claim 1 , wherein the at least one characteristic comprises speaker identification. 
     
     
         3 . The non-transitory computer-readable medium of  claim 1 , wherein at least one characteristic comprises a classification according to a criterion. 
     
     
         4 . The non-transitory computer-readable medium of  claim 3 , wherein the criterion is one of age, gender, mood, and prosody. 
     
     
         5 . The non-transitory computer-readable medium of  claim 1 , wherein the at least one characteristic indicates a weight. 
     
     
         6 . A non-transitory computer-readable medium comprising code effective to cause one or more processors to:
 characterize a speech utterance to determine at least one characteristic;   
       determine from the speech utterance a plurality of phoneme sequence hypotheses;
 tokenize the plurality of phoneme sequence hypotheses to create word sequence hypotheses; 
 prune the plurality of word sequence hypotheses to a set of word sequence hypotheses conditioned on the at least one characteristic; 
 compute probabilities for the word sequence hypotheses from the pruned set of word sequence hypotheses according to a statistical language model; and 
 choose as a transcription a word sequence hypothesis having the highest probability. 
 
     
     
         7 . The non-transitory computer-readable medium of  claim 6 , wherein the at least one characteristic comprises speaker identification. 
     
     
         8 . The non-transitory computer-readable medium of  claim 6 , wherein at least one characteristic comprises a classification according to a criterion. 
     
     
         9 . The non-transitory computer-readable medium of  claim 8 , wherein the criterion is one of age, gender, mood, and prosody. 
     
     
         10 . The non-transitory computer-readable medium of  claim 6 , wherein the at least one characteristic indicates a weight. 
     
     
         11 . A method comprising:
 characterizing a speech utterance to determine at least one characteristic;   determining from the speech utterance a plurality of phoneme sequence hypotheses;   pruning the plurality of phoneme sequence hypotheses to a set of phoneme sequence hypotheses conditioned on the at least one characteristic;   tokenizing the set of phoneme sequence hypotheses to create word sequence hypotheses from the pruned set of phoneme sequence hypotheses;   computing probabilities for the word sequence hypotheses according to a statistical language model; and   choosing as a transcription a word sequence hypothesis having the highest probability.   
     
     
         12 . The method of  claim 11 , wherein the at least one characteristic comprises speaker identification. 
     
     
         13 . The method of  claim 11 , wherein at least one characteristic comprises a classification according to a criterion. 
     
     
         14 . The method of  claim 11 , wherein the criterion is one of age, gender, mood, and prosody. 
     
     
         15 . The method of  claim 11 , wherein the at least one characteristic indicates a weight.

Join the waitlist — get patent alerts

Track US2021335351A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.