Speech Dialog System Aware of Ongoing Conversations
Abstract
Disclosed are systems and methods aware of ongoing conversations and configured to intelligently schedule a speech prompt to an intended addressee. A method for intelligently scheduling a speech prompt in a speech dialog system includes monitoring an acoustic environment to detect an intended addressee's availability for a speech prompt having a measure of urgency corresponding therewith. Based on the intended addressee's availability, the method predicts a time that is convenient to present the speech prompt to the intended addressee, and schedules the speech prompt based on the predicted time and the measure of urgency. A measure of rudeness can be estimated using a cost function that includes cost for presence of an utterance, cost for presence of a conversation, and cost for involvement of the intended addressee in the conversation. Scheduling the speech prompt can include trading off the measure of urgency and the measure of rudeness.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for intelligently scheduling a speech prompt in a speech dialog system, the method comprising:
monitoring an acoustic environment to detect an intended addressee's availability for a speech prompt having a measure of urgency corresponding therewith; based on the intended addressee's availability, predicting a time that is convenient to present the speech prompt to the intended addressee; and scheduling the speech prompt based on the predicted time and the measure of urgency.
2 . The method of claim 1 , wherein monitoring the acoustic environment includes detecting an acoustic signal associated with the acoustic environment to produce a detected acoustic signal, applying speech signal enhancement to the detected acoustic signal to produce an enhanced detected acoustic signal, and generating an enhanced speech signal and a speech activity signal as a function of the enhanced detected acoustic signal.
3 . The method of claim 2 , further comprising detecting dialog from the speech activity signal.
4 . The method of claim 3 , further comprising capturing a video signal associated with the acoustic environment and applying visual speech activity detection to the video signal to generate a visual speech activity signal, wherein the dialog is detected from the speech activity signal and the visual speech activity signal.
5 . The method of claim 3 , further comprising applying voice biometry analysis to the enhanced speech signal to detect involvement of the intended addressee in the dialog.
6 . The method of claim 3 , further comprising:
applying one or more of automatic speech recognition, prosody analysis, and syntactic analysis to the enhanced speech signal to generate one or more speech analysis results; and applying pause prediction to the enhanced speech signal based on the one or more speech analysis results.
7 . The method of claim 6 , wherein predicting the time that is convenient to present the speech prompt includes estimating rudeness of interruption based on the pause prediction and dialog detection to generate a measure of rudeness.
8 . The method of claim 7 , wherein the measure of rudeness is estimated using a cost function that includes cost for presence of an utterance, cost for presence of a conversation, and cost for involvement of the intended addressee in the conversation.
9 . The method of claim 8 , wherein scheduling the speech prompt includes trading off the measure of urgency and the measure of rudeness.
10 . The method of claim 9 , wherein the trading off includes computing an urgency-rudeness ratio as the ratio of the measure of urgency and the measure of rudeness, and wherein the prompt is scheduled based on a comparison of the urgency-rudeness ratio to a threshold.
11 . A speech dialog system for intelligently scheduling a speech prompt, the system comprising:
a dialog manager configured to monitor an acoustic environment to detect an intended addressee's availability for a speech prompt having a measure of urgency corresponding therewith; a scheduler configured to schedule the speech prompt; and a processor in communication with the dialog manager and scheduler, and configured to (i) predict a time that is convenient to present the speech prompt to the intended addressee based on the intended addressee's availability, and (ii) cause the scheduler to schedule the speech prompt based on the predicted time and the measure of urgency.
12 . The system of claim 11 , further comprising:
a microphone system configured to detect an acoustic signal associated with the acoustic environment to produce a detected acoustic signal; and a speech processor in communication with the dialog manager and configured to apply speech signal enhancement to the detected acoustic signal to produce an enhanced detected acoustic signal, the speech processor configured to generate an enhanced speech signal and a speech activity signal as a function of the enhanced detected acoustic signal.
13 . The system of claim 12 , wherein the dialog manager is configured to detect dialog from the speech activity signal.
14 . The system of claim 13 , further comprising:
a camera configured to capture a video signal associated with the acoustic environment; and a video processor in communication with the dialog manager and configured to apply visual speech activity detection to the video signal to generate a visual speech activity signal, wherein the dialog manager is configured to detect the dialog from the speech activity signal and the visual speech activity signal.
15 . The system of claim 13 , further comprising a voice analyzer in communication with the dialog manager and configured to apply voice biometry analysis to the enhanced speech signal to detect involvement of the intended addressee in the dialog.
16 . The system of claim 13 , further comprising a speech recognition engine in communication with the processor and configured to apply one or more of automatic speech recognition, prosody analysis, and syntactic analysis to the enhanced speech signal to generate one or more speech analysis results, wherein the processor is further configured to apply pause prediction to the enhanced speech signal based on the one or more speech analysis results.
17 . The system of claim 16 , wherein the processor is configured to predict the time that is convenient to present the speech prompt by estimating rudeness of interruption based on the pause prediction and dialog detection to generate a measure of rudeness.
18 . The system of claim 17 , wherein the processor is configured to cause the scheduler to schedule the speech prompt by trading off the measure of urgency and the measure of rudeness.
19 . A non-transitory computer-readable medium including computer code instructions stored thereon for intelligently scheduling a speech prompt in a speech dialog system, the computer code instructions, when executed by a processor, cause the system to perform at least the following:
monitor an acoustic environment to detect an intended addressee's availability for a speech prompt having a measure of urgency corresponding therewith; based on the intended addressee's availability, predict a time that is convenient to present the speech prompt to the intended addressee; and schedule the speech prompt based on the predicted time and the measure of urgency.Join the waitlist — get patent alerts
Track US2020349933A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.