US2019279617A1PendingUtilityA1

Voice Characterization-Based Natural Language Filtering

Assignee: SOUNDHOUND INCPriority: Dec 23, 2016Filed: May 23, 2019Published: Sep 12, 2019
Est. expiryDec 23, 2036(~10.4 yrs left)· nominal 20-yr term from priority
Inventors:Karl Stahl
G10L 17/02G10L 15/1822G10L 15/1807G10L 25/63G10L 2015/025G10L 15/19
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An utterance is analyzed to determine a characteristic of the utterance and a transcription hypothesis is generated for the utterance. Grammar rules are then used to parse the transcription hypothesis to produce a plurality of interpretation hypotheses, each having a likelihood score. A set of authorized domains is determined based on the characteristic and the plurality of interpretation hypotheses are filtered according to the set of authorized domains. Of the remaining interpretation hypotheses, one is selected according to their likelihood scores. The characteristic may include one or more characteristics such as mood, prosody, or whether the utterance has a rising intonation.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A non-transitory computer-readable medium comprising code effective to cause one or more processors to:
 characterize a speech utterance to determine at least one characteristic;   recognize the speech utterance, without regard to the at least one characteristic, to produce at least one transcription hypothesis;   parse the at least one transcription hypothesis according to a set of grammar rules to produce a plurality of interpretation hypotheses, each having a corresponding likelihood score;   
       determine a set of authorized domains based on the at least one characteristic; and
 filter the plurality of interpretation hypotheses according to the set of authorized domains; and 
 select a selected interpretation hypothesis from the plurality of interpretation hypotheses according to the likelihood scores thereof. 
 
     
     
         2 . The non-transitory computer-readable medium of  claim 1  wherein the at least one characteristic is mood. 
     
     
         3 . The non-transitory computer-readable medium of  claim 1  wherein the at least one characteristic is prosody. 
     
     
         4 . The non-transitory computer-readable medium of  claim 1  wherein the at least one characteristic is a rising intonation at the end of the speech utterance that indicates a yes or no question. 
     
     
         5 . A system comprising one or more processors and one or more memory devices operably coupled to the one or more processors, the memory devices storing executable code effective to cause one or more processors to:
 characterize a speech utterance to determine at least one characteristic;   recognize the speech utterance, without regard to the at least one characteristic, to produce at least one transcription hypothesis;   parse the at least one transcription hypothesis according to a set of grammar rules to produce a plurality of interpretation hypotheses, each having a corresponding likelihood score;   
       determine a set of authorized domains based on the at least one characteristic; and
 filter the plurality of interpretation hypotheses according to the set of authorized domains; and 
 select a selected interpretation hypothesis from the plurality of interpretation hypotheses according to the likelihood scores thereof. 
 
     
     
         6 . The system of  claim 5  wherein the at least one characteristic is mood. 
     
     
         7 . The system of  claim 5  wherein the at least one characteristic is prosody. 
     
     
         8 . The system of  claim 5  wherein the at least one characteristic is a rising intonation at the end of the speech utterance that indicates a yes or no question. 
     
     
         9 . A method comprising:
 characterizing, by a computer system, a speech utterance to determine at least one characteristic;   recognizing, by the computer system, the speech utterance, without regard to the at least one characteristic, to produce at least one transcription hypothesis;   parsing, by the computer system, the at least one transcription hypothesis according to a set of grammar rules to produce a plurality of interpretation hypotheses, each having a corresponding likelihood score;   
       determining, by the computer system, a set of authorized domains based on the at least one characteristic; and
 filtering, by the computer system, the plurality of interpretation hypotheses according to the set of authorized domains; and 
 selecting, by the computer system, a selected interpretation hypothesis from the plurality of interpretation hypotheses according to the likelihood scores thereof. 
 
     
     
         10 . The method  claim 9  wherein the at least one characteristic is mood. 
     
     
         11 . The method of  claim 9  wherein the at least one characteristic is prosody. 
     
     
         12 . The method of  claim 9  wherein the at least one characteristic is a rising intonation at the end of the speech utterance that indicates a yes or no question.

Join the waitlist — get patent alerts

Track US2019279617A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.