US2022156039A1PendingUtilityA1

Voice Control of Computing Devices

Assignee: AMAZON TECH INCPriority: Dec 8, 2017Filed: Nov 19, 2021Published: May 19, 2022
Est. expiryDec 8, 2037(~11.4 yrs left)· nominal 20-yr term from priority
G10L 15/22G10L 15/1815G06F 3/167G10L 15/30G10L 2015/223G10L 15/16G06F 3/0481G10L 15/142
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for voice control of computing devices are disclosed. Applications may be downloaded and/or accessed by a device having a display, and content associated with the applications may be displayed. Many applications do not allow for voice commands to be utilized to interact with the displayed content. Improvements described herein allow for non-voice-enabled applications to utilize voice commands to interact with displayed content by determining screen data displayed by the device and utilizing the screen data to determine an intent associated with the application. Directive data to perform an action corresponding to the intent may be sent to the device and may be utilized to perform the action on an object associated with the displayed content.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method comprising:
 receiving, from a voice-interface device, audio data representing a user utterance to control output of content;   determining, based at least in part on account data associated with the voice-interface device, devices associated with the voice-interface device;   determining that the content is displayed on a first device of the devices when the audio data is received;   selecting, based at least in part on the content being displayed on the first device when the audio data is received, the first device from the devices to receive a command associated with the user utterance;   identifying, based at least in part on determining that the content is being displayed, metadata associated with the content;   generating the command based at least in part on the metadata; and   sending the command to at least one of the voice-interface device or the first device, the command configured to cause control of the content on the first device.   
     
     
         3 . The method of  claim 2 , wherein:
 the content comprises image data being displayed on the first device;   the first device is a television; and   the voice-interface device is a screenless device.   
     
     
         4 . The method of  claim 2 , further comprising:
 determining that the account data indicates that the first device is linked to the voice-interface device such that the first device is enabled to receive the command; and   wherein selecting the first device comprises selecting the first device based at least in part on the account data indicating that the first device is linked to the voice-interface device.   
     
     
         5 . The method of  claim 2 , further comprising:
 determining, from the account data, that an application associated with the content has been linked to the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the application being linked to the first device.   
     
     
         6 . The method of  claim 2 , further comprising:
 determining, from the audio data, that the user utterance corresponds to a first intent to control output of the content;   determining, from the audio data, that the user utterance corresponds to a second intent to control output of the content, the second intent differing from the first intent;   selecting the first intent instead of the second intent based at least in part on the content being displayed on the first device.   
     
     
         7 . The method of  claim 2 , further comprising:
 determining a previous command sent in response to a previous user utterance received at the voice-interface device;   determining that the previous command was associated with the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the previous command being associated with the first device.   
     
     
         8 . The method of  claim 2 , further comprising:
 determining that the audio data includes a predefined trigger expression associated with the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the audio data including the predefined trigger expression.   
     
     
         9 . The method of  claim 2 , further comprising:
 determining, from historical data indicating previous commands associated with the account data, that an intent associated with the user utterance is historically associated with an action performed by the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the intent being historically associated with the action performed by the first device.   
     
     
         10 . The method of  claim 2 , further comprising:
 determining that an application associated with the content is not enabled for the first device; and   based at least in part on selecting the first device, causing the application to be enabled for the first device.   
     
     
         11 . The method of  claim 2 , further comprising:
 determining that an application associated with the content is not enabled for the first device; and   based at least in part on selecting the first device, causing the voice-interface device to control the content on the first device.   
     
     
         12 . A system comprising:
 one or more processors; and   non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
 receiving, from a voice-interface device, audio data representing a user utterance to control output of content; 
 determining, based at least in part on account data associated with the voice-interface device, devices associated with the voice-interface device; 
 determining that the content is displayed on a first device of the devices when the audio data is received; 
 selecting, based at least in part on the content being displayed on the first device when the audio data is received, the first device from the devices to receive a command associated with the user utterance; 
 generating the command based at least in part on the audio data; and 
 sending the command to at least one of the voice-interface device or the first device, the command configured to cause control of the content on the first device. 
   
     
     
         13 . The system of  claim 12 , wherein:
 the content comprises image data being displayed on the first device;   the first device is a television; and   the voice-interface device is a mobile device.   
     
     
         14 . The system of  claim 12 , the operations further comprising:
 determining that the account data indicates that the first device is enabled in association with the voice-interface device; and   wherein selecting the first device comprises selecting the first device based at least in part on the account data indicating that the first device is enabled in association with the voice-interface device.   
     
     
         15 . The system of  claim 12 , the operations further comprising:
 determining, from the account data, that an application associated with the content has been installed on the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the application being installed on the first device.   
     
     
         16 . The system of  claim 12 , the operations further comprising:
 determining, from the audio data, that the user utterance corresponds to a first intent to control output of the content;   determining, from the audio data, that the user utterance corresponds to a second intent to control output of the content, the second intent differing from the first intent;   selecting the first intent to utilize for generation of the command based at least in part on the content being displayed on the first device.   
     
     
         17 . The system of  claim 12 , the operations further comprising:
 determining a previous command sent in response to a previous user utterance received in association with the account data;   determining that the previous command was associated with the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the previous command being associated with the first device.   
     
     
         18 . The system of  claim 12 , the operations further comprising:
 determining that the audio data includes a predefined word associated with the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the audio data including the predefined trigger expression.   
     
     
         19 . The system of  claim 12 , the operations further comprising:
 determining, from historical data associated with the account data, that an intent associated with the user utterance is historically associated with an action performed by the first device; and   wherein selecting the first device comprises selecting the first device based at least in part on the intent being historically associated with the action performed by the first device.   
     
     
         20 . The system of  claim 12 , the operations further comprising:
 determining that an application associated with the content is not installed on the first device; and   based at least in part on selecting the first device, causing the application to be installed on the first device.   
     
     
         21 . The system of  claim 12 , the operations further comprising:
 determining that an application associated with the content is not installed on the first device; and   based at least in part on application not being installed on the first device, causing the voice-interface device to control the content on the first device.

Join the waitlist — get patent alerts

Track US2022156039A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.