US2017116990A1PendingUtilityA1

Visual confirmation for a recognized voice-initiated action

Assignee: GOOGLE INCPriority: Jul 31, 2013Filed: Jan 5, 2017Published: Apr 27, 2017
Est. expiryJul 31, 2033(~7 yrs left)· nominal 20-yr term from priority
G06F 3/04817G10L 15/22G10L 2015/223G06F 3/167G10L 15/1815G01C 21/3608G10L 2015/228
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Techniques described herein provide a computing device configured to provide an indication that the computing device has recognized a voice-initiated action. In one example, a method is provided for outputting, by a computing device and for display, a speech recognition graphical user interface (GUI) having at least one element in a first visual format. The method further includes receiving, by the computing device, audio data and determining, by the computing device, a voice-initiated action based on the audio data. The method also includes outputting, while receiving additional audio data and prior to executing a voice-initiated action based on the audio data, and for display, an updated speech recognition GUI in which the at least one element is displayed in a second visual format, different from the first visual format, to indicate that the voice-initiated action has been identified.

Claims

exact text as granted — not AI-modified
1 . A method comprising:
 displaying, by a computing device, a speech recognition graphical user interface (GUI) including a non-textual element that is displayed in an initial visual format that indicates the computing device is executing in speech recognition mode;   responsive to determining, based on first audio data of a voice command, a first voice-initiated action from a plurality of voice-initiated actions, while receiving second audio data of the voice command, and prior to performing the voice command, displaying the non-textual element in a first visual format that corresponds to the first voice-initiated action, wherein the first visual format is different from the initial visual format;   after receiving the second audio data of the voice command, determining, based on the second audio data, a second voice-initiated action from the plurality of voice-initiated actions that is associated with the first audio data of the voice command, wherein the second voice-initiated action is different than the first voice-initiated action;   responsive to determining the second voice-initiated action, while receiving third audio data of the voice command, and prior to performing the voice command, displaying the non-textual element in a third visual format that corresponds to the second voice-initiated action, wherein the third visual format is different from the first and second visual formats; and   after receiving the third audio data of the voice command, executing, by the computing device, based on the first, second, and third audio data, an application that performs the second voice-initiated action.   
     
     
         2 . The method of  claim 1 , wherein:
 the application is a second application executing at the computing device; and   the first voice-initiated action is associated with a first application executing at the computing device that is different than the second application.   
     
     
         3 . The method of  claim 2 , wherein the computing device executes in speech recognition mode to display the speech recognition GUI by executing a third application that is different than the first and second applications. 
     
     
         4 . The method of  claim 1 , wherein each voice-initiated action from the plurality of voice-initiated actions corresponds to different visual format of the non-textual element. 
     
     
         5 . The method of  claim 1 , further comprising:
 determining, based on the first audio data of the voice command, one or more words of the voice command; and   determining, based on the one or more words of the voice command, the first voice-initiated action.   
     
     
         6 . The method of  claim 5 , wherein determining the second voice-initiated action comprises:
 determining, based on the second audio data of the voice command, a new meaning for the one or more words of the voice command; and   determining, based on the new meaning, the second voice-initiated action.   
     
     
         7 . The method of  claim 5 , wherein determining the first voice-initiated action comprises determining the first voice-initiated action based at least partially on a comparison of at least one of the one or more words of the voice command to at least one respective word associated with each voice-initiated action from the plurality of voice-initiated actions. 
     
     
         8 . The method of  claim 7 , wherein the at least one respective word associated with each voice-initiated action from the plurality of voice-initiated actions comprises a respective verb corresponding to that voice-initiated action. 
     
     
         9 . The method of  claim 1 , further comprising
 determining, a context based at least part on data from the computing device; and   determining, based at least partially on the context and the first audio data, the first voice-initiated action.   
     
     
         10 . A computing device comprising:
 a display device;   a microphone;   one or more processors; and   a memory storing instructions that, when executed, cause the one or more processors to:   display, at the display device, a speech recognition graphical user interface (GUI) including a non-textual element that is displayed in an initial visual format that indicates the computing device is executing in speech recognition mode;   responsive to determining, based on first audio data of a voice command received by the microphone, a first voice-initiated action from a plurality of voice-initiated actions, while the microphone receives second audio data of the voice command, and prior to the one or more processors performing the voice command, display, at the display device, the non-textual element in a first visual format that corresponds to the first voice-initiated action, wherein the first visual format is different from the initial visual format;   after the microphone receives the second audio data of the voice command, determine, based on the second audio data, a second voice-initiated action from the plurality of voice-initiated actions that is associated with the first audio data of the voice command, wherein the second voice-initiated action is different than the first voice-initiated action;   responsive to determining the second voice-initiated action, while the microphone receives third audio data of the voice command, and prior to the one or more processors performing the voice command, display, at the display device, the non-textual element in a third visual format that corresponds to the second voice-initiated action, wherein the third visual format is different from the first and second visual formats; and   after the microphone receives the third audio data of the voice command, execute, based on the first, second, and third audio data, an application that performs the second voice-initiated action.   
     
     
         11 . The computing device of  claim 10 , wherein each voice-initiated action from the plurality of voice-initiated actions corresponds to different visual format of the non-textual element. 
     
     
         12 . The computing device of  claim 10 , wherein the instructions, when executed, further cause the one or more processors to:
 determine, based on the first audio data of the voice command, one or more words of the voice command; and   determine, based on the one or more words of the voice command, the first voice-initiated action.   
     
     
         13 . The computing device of  claim 12 , wherein the instructions, when executed, further cause the one or more processors to determine the second voice-initiated action by:
 determining, based on the second audio data of the voice command, a new meaning for the one or more words of the voice command; and   determining, based on the new meaning, the second voice-initiated action.   
     
     
         14 . The computing device of  claim 12 , wherein the instructions, when executed, further cause the one or more processors to determine the first voice-initiated action by determining the first voice-initiated action based at least partially on a comparison of at least one of the one or more words of the voice command to at least one respective word associated with each voice-initiated action from the plurality of voice-initiated actions. 
     
     
         15 . The computing device of  claim 14 , wherein the at least one respective word associated with each voice-initiated action from the plurality of voice-initiated actions comprises a respective verb corresponding to that voice-initiated action. 
     
     
         16 . The computing device of  claim 10 , wherein the instructions, when executed, further cause the one or more processors to:
 determine a context based at least part on data from the computing device; and   determine, based at least partially on the context and the first audio data, the first voice-initiated action.   
     
     
         17 . A computer-readable storage medium encoded with instructions that, when executed by one or more processors of a computing device, cause the one or more processors to:
 display a speech recognition graphical user interface (GUI) including a non-textual element that is displayed in an initial visual format that indicates the computing device is executing in speech recognition mode;   responsive to determining, based on first audio data of a voice command, a first voice-initiated action from a plurality of voice-initiated actions, while receiving second audio data of the voice command, and prior to performing the voice command, display the non-textual element in a first visual format that corresponds to the first voice-initiated action, wherein the first visual format is different from the initial visual format;   after the microphone receives the second audio data of the voice command, determine, based on the second audio data, a second voice-initiated action from the plurality of voice-initiated actions that is associated with the first audio data of the voice command, wherein the second voice-initiated action is different than the first voice-initiated action;   responsive to determining the second voice-initiated action, while the microphone receives third audio data of the voice command, and prior to the one or more processors performing the voice command, display, at the display device, the non-textual element in a third visual format that corresponds to the second voice-initiated action, wherein the third visual format is different from the first and second visual formats; and   after receiving the third audio data of the voice command, execute, based on the first, second, and third audio data, an application that performs the second voice-initiated action.   
     
     
         18 . The computer-readable storage medium of  claim 17 , wherein:
 the application is a second application executing at the computing device; and   the first voice-initiated action is associated with a first application executing at the computing device that is different than the second application.   
     
     
         19 . The computer-readable storage medium of  claim 18 , wherein the computing device executes in speech recognition mode to display the speech recognition GUI by executing a third application that is different than the first and second applications 
     
     
         20 . The computer-readable storage medium of  claim 17 , wherein each voice-initiated action from the plurality of voice-initiated actions corresponds to different visual format of the non-textual element.

Join the waitlist — get patent alerts

Track US2017116990A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.