US2017249935A1PendingUtilityA1

System and method for estimating the reliability of alternate speech recognition hypotheses in real time

Assignee: NUANCE COMMUNICATIONS INCPriority: Oct 23, 2009Filed: May 15, 2017Published: Aug 31, 2017
Est. expiryOct 23, 2029(~3.2 yrs left)· nominal 20-yr term from priority
G10L 15/08G10L 15/01G10L 15/083G10L 15/28G10L 15/14G10L 15/22G10L 15/04
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Disclosed herein are systems, methods, and computer-readable storage media for estimating reliability of alternate speech recognition hypotheses. A system configured to practice the method receives an N-best list of speech recognition hypotheses and features describing the N-best list, determines a first probability of correctness for each hypothesis in the N-best list based on the received features, determines a second probability that the N-best list does not contain a correct hypothesis, and uses the first probability and the second probability in a spoken dialog. The features can describe properties of at least one of a lattice, a word confusion network, and a garbage model. In one aspect, the N-best lists are not reordered according to reranking scores. The determination of the first probability of correctness can include a first stage of training a probabilistic model and a second stage of distributing mass over items in a tail of the N-best list.

Claims

exact text as granted — not AI-modified
We claim: 
     
         1 . A method comprising:
 receiving an N-best list of speech recognition hypotheses from a speech utterance;   determining, via a processor and based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and   using the probability of correctness in a spoken dialog.   
     
     
         2 . The method of  claim 1 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network. 
     
     
         3 . The method of  claim 1 , wherein the processor is configured to perform speech language generation. 
     
     
         4 . The method of  claim 1 , wherein determining the probability of correctness is performed in two stages. 
     
     
         5 . The method of  claim 4 , wherein a first stage of the two stages comprises training a discriminative model P a . 
     
     
         6 . The method of  claim 5 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses. 
     
     
         7 . A system comprising:
 a processor; and   a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising:
 receiving an N-best list of speech recognition hypotheses from a speech utterance; 
 determining, based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and 
 using the probability of correctness in a spoken dialog. 
   
     
     
         8 . The system of  claim 7 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network. 
     
     
         9 . The system of  claim 7 , wherein the processor is configured to perform speech language generation. 
     
     
         10 . The system of  claim 7 , wherein determining the probability of correctness is performed in two stages. 
     
     
         11 . The system of  claim 10 , wherein a first stage of the two stages comprises training a discriminative model P a . 
     
     
         12 . The system of  claim 11 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses. 
     
     
         13 . A non-transitory computer-readable storage medium having instructions stored which, when executed by a computing device, cause the computing device perform operations comprising:
 receiving an N-best list of speech recognition hypotheses from a speech utterance;   determining, via a processor and based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and   using the probability of correctness in a spoken dialog.   
     
     
         14 . The non-transitory computer-readable storage medium of  claim 13 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network. 
     
     
         15 . The non-transitory computer-readable storage medium of  claim 13 , wherein the processor is configured to perform speech language generation. 
     
     
         16 . The non-transitory computer-readable storage medium of  claim 13 , wherein determining the probability of correctness is performed in two stages. 
     
     
         17 . The non-transitory computer-readable storage medium of  claim 16 , wherein a first stage of the two stages comprises training a discriminative model P a . 
     
     
         18 . The non-transitory computer-readable storage medium of  claim 17 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses.

Join the waitlist — get patent alerts

Track US2017249935A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.