System and method for a language understanding conversational system
Abstract
A virtual assistant device recognizes multiple wake-up phrases. In response to a particular wake-up phrase the device sends speech audio to either a default or a third party virtual assistant server. A virtual assistant server can receive speech audio and an indication of which of multiple wake-up phrases was used and, accordingly, send the speech audio, or text recognized from the speech audio using automatic speech recognition, to a third party server. A response from the third party server can be voice audio or text for the virtual assistant server to synthesize distinctively corresponding to the wake-up phrase.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A service for voice conversational interfaces comprising:
receiving speech audio content indicating a wake-up phrase that identifies a knowledge domain; performing speech recognition; performing natural language interpretation, wherein the natural language interpretation uses the knowledge domain to determine a user intent; taking a fulfilment action in response to the user intent; and sending a response message in response to the fulfillment action.
2 . The service of claim 1 , wherein the speech audio content is received through a runtime API.
3 . The service of claim 1 , wherein the speech audio content comprises a user ID.
4 . The service of claim 1 , wherein actions occur through repeated conversation turns in a session and the speech audio content comprises a session ID.
5 . The service of claim 1 , wherein the fulfilment action is to conduct a dialog.
6 . The service of claim 1 , wherein the fulfilment action is to request necessary slot data and add it to conversation state history.
7 . The service of claim 1 wherein the fulfilment action provides lifelike conversational interactions.
8 . A language understanding bot framework comprising:
a plurality of keyword recognition models that may be invoked simultaneously in a device; a plurality of language understanding models, each representing a knowledge domain; a configurable system of services enabled to receive a keyword recognition result that indicates which of the plurality of recognition models is activated and conditionally invokes a corresponding language understanding recognizer.
9 . The framework of claim 8 further comprising runtime language generation that generates a response to a user according to the language understanding recognizer result, wherein the language understanding recognizer produces a result.
10 . The framework of claim 8 further comprising a speech recognition service, which follows the keyword recognition result, that converts speech audio to text and provides the text to the language understanding recognizer.
11 . A non-transitory computer readable medium for storing code that is executed by the processor to cause a system to:
receive speech audio content that includes a wake-up phrase identify a knowledge domain based on the wake-up phrase; perform speech recognition; perform natural language interpretation, wherein the natural language interpretation uses the knowledge domain to determine a user's intent; and take a fulfilment action in response to the user intent.
12 . The non-transitory computer readable medium of claim 11 , wherein the processor executes code and causes the system to send a response message in response to the fulfillment action.Join the waitlist — get patent alerts
Track US2020410983A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.