US2019164553A1PendingUtilityA1

Speech-enabled system with domain disambiguation

Assignee: SOUNDHOUND INCPriority: Mar 10, 2017Filed: Jan 10, 2019Published: May 30, 2019
Est. expiryMar 10, 2037(~10.6 yrs left)· nominal 20-yr term from priority
Inventors:Rainer Leeb
G10L 15/18G10L 15/1815G10L 2015/221G10L 15/22G06F 40/30G10L 2015/225G10L 15/02G06F 17/2785
38
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems perform methods of interpreting spoken utterances from a user and responding to the utterances by providing requested information or performing a requested action. The utterances are interpreted in the context of multiple domains. Each interpretation is assigned a relevancy score based on how well the interpretation represents what the speaker intended. Interpretations having a relevancy score below a threshold for its associated domain are discarded. A remaining interpretation is chosen based on choosing the most relevant domain for the utterance. The user may be prompted to provide disambiguation information that can be used to choose the best domain. Storing past associations of utterance representation and domain choice allows for measuring the strength of correlation between uttered words and phrases with relevant domains. This correlation strength information may allow the system to automatically disambiguate alternate interpretations without requiring user input.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of disambiguating natural language utterances, the method comprising:
 using at least one computer to:
 interpret a natural language utterance according to a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith; 
 interpret the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area and comprises a different set of grammar rules; 
 
 determine whether the first relevancy score exceeds the first relevancy threshold; 
 determine whether the second relevancy score exceeds the second relevancy threshold; 
 present, to a user, a list of candidate domains that includes (i) the first domain if the first relevancy score is determined to exceed the first relevancy threshold and (ii) the second domain if the second relevancy score is determined to exceed the second relevancy threshold; 
 ask the user to choose a domain from the presented list; and 
 present, to the user, only one of the first interpretation and the second interpretation according to a domain chosen by the user. 
   
     
     
         2 . The method of  claim 1 , further comprising using the at least one computer to increment a value of a counter representing the chosen domain. 
     
     
         3 . The method of  claim 2 , wherein the candidate domains are presented to the user in an order based on the value of the counter. 
     
     
         4 . The method of  claim 1 , further comprising using the at least one computer to store an indication of the most recently chosen domain. 
     
     
         5 . The method of  claim 4  wherein the candidate domains are presented to the user in an order based on the indication of the most recently chosen domain. 
     
     
         6 . The method of  claim 1  further comprising:
 using the at least one computer to:
 store a record in a database, the record comprising:
 a representation of the natural language utterance; and 
 the choice of domain for the natural language utterance. 
 
 
 
     
     
         7 . The method of  claim 1  further comprising:
 using the at least one computer to:
 store a record in a database, the record comprising:
 the interpretation of the natural language utterance according to the chosen domain; and 
 the choice of domain. 
 
 
 
     
     
         8 . The method of  claim 1 , wherein a particular threshold is assigned to each domain and at least two particular thresholds are different for at least two domains. 
     
     
         9 . The method of  claim 1 , wherein audio speech is used to (i) present the list of candidate domains to the user and (ii) ask the user to choose the domain from the presented list. 
     
     
         10 . A method of disambiguating natural language utterances, the method comprising:
 using at least one computer to:
 interpret a natural language utterance according a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith; 
 interpret the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area and comprises a different set of grammar rules; 
 
 determine whether the first relevancy score exceeds the first relevancy threshold; 
 determine whether the second relevancy score exceeds the second relevancy threshold; 
 determine a number of candidate domains, including the first domain and the second domain, for which the corresponding relevancy score exceeds the relevancy threshold associated therewith; and 
 responsive to the number of candidate domains being greater than a maximum number of domains that is reasonable to present to a user for disambiguation, ask the user to provide a general clarification; 
 receive a response utterance with clarification information to determine a domain; and 
 present, to the user, only one of the first interpretation and the second interpretation according to determined domain. 
   
     
     
         11 . The method of  claim 10  wherein the maximum number of domains that is reasonable to present to the user is based on environment information. 
     
     
         12 . A non-transitory computer readable medium storing code that, when executed by at least one computer, causes the one or more computer to:
 interpret a natural language utterance according to a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith;   interpret the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area and comprises a different set of grammar rules; 
   determine whether the first relevancy score exceeds the first relevancy threshold;   determine whether the second relevancy score exceeds the second relevancy threshold;   present, to a user, a list of candidate domains that includes (i) the first domain if the first relevancy score is determined to exceed the first relevancy threshold and (ii) the second domain if the second relevancy score is determined to exceed the second relevancy threshold;   ask the user to choose a domain from the presented list; and   present, to the user, only one of the first interpretation and the second interpretation according to a domain chosen by the user.   
     
     
         13 . A speech-enabled system for disambiguating natural language utterances, the speech-enabled system comprising:
 means for performing disambiguation by:
 interpret a natural language utterance according to a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith; 
 interpret the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area and comprises a different set of grammar rules; 
 
 determine whether the first relevancy score exceeds the first relevancy threshold; 
 determine whether the second relevancy score exceeds the second relevancy threshold; 
 present, to a user, a list of candidate domains that includes (i) the first domain if the first relevancy score is determined to exceed the first relevancy threshold and (ii) the second domain if the second relevancy score is determined to exceed the second relevancy threshold; 
 ask the user to choose a domain from the presented list; and 
 present, to the user, only one of the first interpretation and the second interpretation according to a domain chosen by the user. 
   
     
     
         14 . An automotive platform for disambiguating natural language utterances, the automotive platform comprising:
 a speech capture module enabled to capture a spoken utterance from a user;   a speech recognition module that:
 interprets a natural language utterance according to a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith; and 
 interprets the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area and comprises a different set of grammar rules; 
 
 determine whether the first relevancy score exceeds the first relevancy threshold; 
 determine whether the second relevancy score exceeds the second relevancy threshold; 
   a speech generation module enabled to:
 produce speech, to a user, that comprises a list of candidate domains that includes (i) the first domain if the first relevancy score is determined to exceed the first relevancy threshold and (ii) the second domain if the second relevancy score is determined to exceed the second relevancy threshold; 
 ask the user to choose a domain from the list; and 
 produce speech, to the user, that comprises only one of the first interpretation and the second interpretation according to a domain chosen by the user. 
   
     
     
         15 . A method of disambiguating natural language utterances, the method comprising:
 using at least one computer to:
 interpret a natural language utterance according to a first domain to create (i) a first interpretation that is unique to the first domain and (ii) a corresponding first relevancy score, the first domain having a first relevancy threshold associated therewith; 
 interpret the same natural language utterance according to a second domain to create (i) a second interpretation that is unique to the second domain and different from the first interpretation and (ii) a corresponding second relevancy score, the second domain having a second relevancy threshold associated therewith,
 wherein each of the first and second domains represents a different subject area for interpretation; 
 
 present, to a user, a list of candidate domains that includes (i) the first domain if the first relevancy score has been determined to exceed the first relevancy threshold and (ii) the second domain if the second relevancy score has been determined to exceed the second relevancy threshold; 
 ask the user to choose a domain from the presented list; and 
 present, to the user, only one of the first interpretation and the second interpretation according to a domain chosen by the user.

Join the waitlist — get patent alerts

Track US2019164553A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.