Using voice input to control a user interface within an application
Abstract
Techniques include a method of providing a user interface on a device. The user interface has at least a first display portion, a second display portion and a third display portion, the second display portion including a link to the third display portion that, when activated by a user, cause the device to present the third display portion of the user interface, the first display portion not including the link to the third display portion. The method includes causing the first display portion to be displayed. The method further includes receiving audible input while the first display portion is being displayed. The method further includes determining that the third display portion corresponds to an utterance in the audible input based at least in part on labels determined to match the utterance. The method further includes causing the third display portion to be displayed.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A system, comprising:
a processor;
a sensor operably connected to the processor;
a display operably connected to the processor; and
one or more computer-readable media storing instructions which, when executed by processor, cause the processor to:
display, via the display, a first portion of a user interface without a second portion of the user interface different from the first portion;
receive, from the sensor and while the first portion is being displayed, an audible input;
determine an utterance included in the audible input;
determine, by providing the utterance, as an input, to a machine learning (ML) model, a label associated with the utterance;
identify, based on the label, the second portion of the user interface; and
display, via the display, the second portion of the user interface without the first portion.
2. The system of claim 1 , wherein the label is indicative of a particular functionality of a set of functionalities provided by the user interface.
3. The system of claim 1 , wherein the ML model is trained on a dataset comprising a set of training utterances, each training utterance corresponding to a label of a set of labels.
4. The system of claim 3 , wherein the audible input is a first audible input, and the instructions further cause the processor to:
receive, from the sensor, a second audible input;
determine, based on an output of the ML model, that the second audible input corresponds to an unknown label; and
provide, based on determining that the second audible input corresponds to an unknown label, the second audible input to be added to the dataset.
5. The system of claim 1 , wherein the system further comprises a speaker operably connected to the processor, the audible input is a first audible input, and the instructions further cause the processor to:
provide, via the speaker, an audio indication of the second portion; and
receive, from the sensor and in response to the audio indication, a second audible input.
6. The system of claim 5 , wherein the instructions further cause the processor to:
determine that the second audible input is indicative of a confirmation,
wherein the second portion is displayed without the first portion based on the confirmation.
7. The system of claim 5 , wherein the instructions further cause the processor to:
determine that the second audible input is indicative of a rejection;
display, based on determining that the second audible input is indicative of the rejection, the first portion without the second portion; and
provide the first audible input as training data for additional training of the ML model.
8. The system of claim 1 , wherein the instructions further cause the processor to:
determine, based on an output of the ML model, a confidence level associated with the label, the confidence level being indicative of a degree to which the label corresponds to the utterance; and
determine, that the confidence level is higher than a threshold,
wherein the second portion is displayed without the first portion based on determining that the confidence level is higher than the threshold.
9. The system of claim 1 , wherein the audible input is a first audible input, and the instructions further cause the processor to:
receive, from the sensor, a second audible input;
determine that a value of a similarity metric between a first representation of the first audible input and a second representation of the second audible input is higher than a threshold; and
adding, based on determining that the value of the similarity is higher than the threshold, the first representation and the second representation to a dataset for training the ML model.
10. The system of claim 1 , wherein the instructions further cause the processor to:
receive data associated with a user of the user interface,
the processor identifying the second portion based on the data associated with the user.
11. A method, comprising:
displaying, by a processor and via a display operably connected to the processor, a first portion of a user interface without a second portion of the user interface different from the first portion;
receiving, by the processor and while the first portion is being displayed, an audible input;
determining, by the processor, an utterance included in the audible input;
determining, by the processor and based on inputting the utterance to a trained machine learning (ML) model, a label associated with the utterance;
identifying, by the processor and based on the label, the second portion of the user interface; and
displaying, by the processor and via the display, the second portion of the user interface without the first portion.
12. The method of claim 11 , further comprising:
receiving, by the processor, data associated with a user of the user interface,
the processor identifying the second portion based on the data associated with the user.
13. The method of claim 11 , wherein the ML model is trained on a dataset
comprising utterances and corresponding labels, the method further comprising:
adding, by the processor, the utterance and the label to the dataset for additional training of the ML model.
14. The method of claim 11 , further comprising:
providing, by the processor, an audio indication or a haptic indication indicative of the second portion; and
receiving, by the processor and in response to the audio indication or the haptic indication, a confirmation from a user of the user interface,
the processor displaying the second portion based on receiving the confirmation.
15. The method of claim 11 , further comprising:
determining, by the processor, a confidence level associated with the label; and
determining, by the processor, that the confidence level is higher than a threshold,
the processor displaying the second portion based at least in part on determining that the confidence level is higher than the threshold.
16. The method of claim 11 , wherein:
the label is indicative of a particular functionality of a set of functionalities provided by the user interface, and
the second portion of the user interface is associated with the particular functionality.
17. The method of claim 16 , wherein:
the particular functionality is different from a first functionality of the first portion of the user interface, and
the first portion of the user interface does not include a link to the second portion.
18. A system, comprising:
a means for displaying a user interface;
a means for receiving audible input;
a means for providing audible output;
a means for storing executable instructions; and
a means for executing the executable instructions, the means for executing being configured to:
display, via the means for displaying a user interface, a first portion of a user interface associated with a first functionality;
receive, via the means for receiving audible input, a first audible input;
determine, based on first the audible input, an utterance included in the first audible input;
determine, by inputting the utterance to a trained machine learning (ML) model, a label associated with the utterance;
identify, based on the label, a second portion of the user interface associated with a second functionality, different from the first functionality;
provide, via the means for providing audible output, an audible indication of the second portion;
receive, via the means for receiving audible input, a second audible input in response to the indication; and
display, via the means for displaying, and based on the label and the indication, the second portion of the user interface.
19. The system of claim 18 , further comprising a means for electronic communication via a communication network, wherein the means of executing is further configured to:
receive, via the means for electronic communication, data associated with a user of the user interface,
wherein identifying the second portion is further based on the data associated with the user.
20. The system of claim 18 , wherein the first functionality or the second functionality comprises one of: promoting a product, filing an insurance claim, or performing a banking transaction.Join the waitlist — get patent alerts
Track US12045543B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.