System and method for estimating the reliability of alternate speech recognition hypotheses in real time
Abstract
Disclosed herein are systems, methods, and computer-readable storage media for estimating reliability of alternate speech recognition hypotheses. A system configured to practice the method receives an N-best list of speech recognition hypotheses and features describing the N-best list, determines a first probability of correctness for each hypothesis in the N-best list based on the received features, determines a second probability that the N-best list does not contain a correct hypothesis, and uses the first probability and the second probability in a spoken dialog. The features can describe properties of at least one of a lattice, a word confusion network, and a garbage model. In one aspect, the N-best lists are not reordered according to reranking scores. The determination of the first probability of correctness can include a first stage of training a probabilistic model and a second stage of distributing mass over items in a tail of the N-best list.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
receiving an N-best list of speech recognition hypotheses from a speech utterance; determining, via a processor and based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and using the probability of correctness in a spoken dialog.
2 . The method of claim 1 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network.
3 . The method of claim 1 , wherein the processor is configured to perform speech language generation.
4 . The method of claim 1 , wherein determining the probability of correctness is performed in two stages.
5 . The method of claim 4 , wherein a first stage of the two stages comprises training a discriminative model P a .
6 . The method of claim 5 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses.
7 . A system comprising:
a processor; and a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising:
receiving an N-best list of speech recognition hypotheses from a speech utterance;
determining, based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and
using the probability of correctness in a spoken dialog.
8 . The system of claim 7 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network.
9 . The system of claim 7 , wherein the processor is configured to perform speech language generation.
10 . The system of claim 7 , wherein determining the probability of correctness is performed in two stages.
11 . The system of claim 10 , wherein a first stage of the two stages comprises training a discriminative model P a .
12 . The system of claim 11 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses.
13 . A non-transitory computer-readable storage medium having instructions stored which, when executed by a computing device, cause the computing device perform operations comprising:
receiving an N-best list of speech recognition hypotheses from a speech utterance; determining, via a processor and based on a feature set evaluated by an algorithm, a probability of correctness for each hypothesis in the N-best list of speech recognition hypotheses, the feature set comprising a count of words, an acoustic score, and an indication of problematic words; and using the probability of correctness in a spoken dialog.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein the N-best list of speech recognition hypotheses are stored in a word confusion network.
15 . The non-transitory computer-readable storage medium of claim 13 , wherein the processor is configured to perform speech language generation.
16 . The non-transitory computer-readable storage medium of claim 13 , wherein determining the probability of correctness is performed in two stages.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein a first stage of the two stages comprises training a discriminative model P a .
18 . The non-transitory computer-readable storage medium of claim 17 , wherein a second stage of the two stages comprises distributing mass over items in a tail of the N-best list of speech recognition hypotheses.Join the waitlist — get patent alerts
Track US2017249935A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.