Interface for a virtual digital assistant
Abstract
The digital assistant displays a digital assistant object in an object region of a display screen. The digital assistant then obtains at least one information item based on a speech input from a user. Upon determining that the at least one information item can be displayed in its entirety in the display region of the display screen, the digital assistant displays the at least one information item in the display region, where the display region and the object region are not visually distinguishable from one another. Upon determining that the at least one information item cannot be displayed in its entirety in the display region of the video display screen, the digital assistant displays a portion of the at least one information item in the display region, where the display region and the object region are visually distinguishable from one another.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . An electronic device, comprising:
one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
receiving spoken user input via an input device;
generating a first plurality of candidate interpretations of the received spoken user input;
generating a second plurality of candidate interpretations, wherein the second plurality of candidate interpretations is a subset of the first plurality of candidate interpretations;
deriving a representation of user intent based on the second plurality of candidate interpretations, wherein an association between the user intent and a respective candidate interpretation is added to context information;
identifying, based on the user intent, at least one task;
executing the at least one task; and
providing an output based on the at least one executed task.
3 . The device of claim 2 , the one or more programs including instructions for:
prompting the user via a conversational interface; and receiving the spoken user input via the conversational interface; and converting the spoken user input to a text representation.
4 . The device of claim 3 , wherein converting the spoken user input to a text representation comprises:
generating a plurality of candidate text interpretations of the spoken user input; and ranking at least a subset of the generated candidate text interpretations; and wherein at least one of the generating and ranking steps is performed using the context information.
5 . The device of claim 4 , wherein the context information used in at least one of the generating and ranking comprises at least one selected from the group consisting of:
acoustic environment data describing an acoustic environment in which the spoken user input is received; data received from at least one sensor; vocabulary obtained from a database associated with the user; vocabulary associated with application preferences; vocabulary obtained from usage history; and current dialog state.
6 . The device of claim 2 , the one or more programs including instructions for:
prompting the user by generating at least one prompt based at least in part on the context information.
7 . The device of claim 2 , the one or more programs including instructions for:
disambiguating the received spoken user input based on acoustic environment data of received context information to derive a representation of user intent by performing natural language processing on the received spoken user input based at least in part on the context information.
8 . The device of claim 7 , wherein the context information comprises at least one selected from the group consisting of:
data describing an event; application context; input previously provided by the user; known information about the user; location; date; environmental conditions; and history.
9 . The device of claim 2 , the one or more programs including instructions for:
identifying at least one task and at least one parameter for the task by identifying at least one task and at least one parameter for the task based at least in part on the context information.
10 . The device of claim 9 , wherein the context information used in identifying at least one task and at least one parameter for the task comprises at least one selected from the group consisting of:
data describing an event; data from a database associated with the user; data received from at least one sensor; application context; input previously provided by the user; known information about the user; location; date; environmental conditions; and history.
11 . The device of claim 2 , the one or more programs including instructions for:
generating a dialog response based at least in part on the context information.
12 . The device of claim 11 , wherein the context information used in generating a dialog response comprises at least one selected from the group consisting of:
data from a database associated with the user; application context; input previously provided by the user; known information about the user; location; date; environmental conditions; and history.
13 . The device of claim 2 , wherein the context information comprises at least one selected from the group consisting of:
context information stored at a server; and context information stored at a client.
14 . The device of claim 2 , the one or more programs including instructions for:
receiving the context information from a context source by:
requesting the context information from a context source; and
receiving the context information in response to the request.
15 . The device of claim 2 , the one or more programs including instructions for:
receiving context information from a context source by:
receiving at least a portion of the context information prior to receiving the spoken user input.
16 . The device of claim 2 , the one or more programs including instructions for:
receiving at least a portion of the context information after receiving the spoken user input.
17 . The device of claim 2 , the one or more programs including instructions for:
receiving static context information as part of an initialization step; and receiving additional context information after receiving the spoken user input.
18 . The device of claim 2 , the one or more programs including instructions for:
receiving push notification of a change in context information; and responsive to the push notification, updating locally stored context information.
19 . The device of claim 2 , wherein the electronic device corresponds to at least one of:
a telephone; a smartphone; a tablet computer; a laptop computer; a personal digital assistant; a desktop computer; a kiosk; a consumer electronic device; a consumer entertainment device; a music player; a camera; a television; an electronic gaming unit; and a set-top box.
20 . A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first electronic device, the one or more programs including instructions for:
receiving spoken user input via an input device; generating a first plurality of candidate interpretations of the received spoken user input; generating a second plurality of candidate interpretations, wherein the second plurality of candidate interpretations is a subset of the first plurality of candidate interpretations; deriving a representation of user intent based on the second plurality of candidate interpretations, wherein an association between the user intent and a respective candidate interpretation is added to context information; identifying, based on the user intent, at least one task; executing the at least one task; and providing an output based on the at least one executed task.
21 . A computer-implemented method, comprising:
at an electronic device with one or more processors and memory:
receiving spoken user input via an input device;
generating a first plurality of candidate interpretations of the received spoken user input;
generating a second plurality of candidate interpretations, wherein the second plurality of candidate interpretations is a subset of the first plurality of candidate interpretations;
deriving a representation of user intent based on the second plurality of candidate interpretations, wherein an association between the user intent and a respective candidate interpretation is added to context information;
identifying, based on the user intent, at least one task;
executing the at least one task; and
providing an output based on the at least one executed task.Join the waitlist — get patent alerts
Track US2023409283A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.