Implicit target selection for multiple audio playback devices in an environment
Abstract
A user can utter a voice command in an environment where multiple audio playback devices are located to have audio output on a single device, or a predefined group of devices in a synchronized manner. In instances when the voice command uttered by the user does not specify a target for audio output, an implicit target selection algorithm can evaluate one or more criteria to determine an appropriate target for output of the audio corresponding to the voice command. An example criterion is met if a predetermined time period has lapsed since a last utterance was detected by a device in the environment. However, other criteria can be evaluated for determining a target output device(s).
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
determining, based at least in part on first data representing a first utterance, first audio content and one or more output devices; sending second data for output of the first audio content by the one or more output devices; determining, based at least in part on third data representing a second utterance, second audio content; determining that the third data omits a specific output device; determining, based at least in part on the determining that the third data omits the specific output device, a stored preference specifying one or more preferred output devices; selecting, based at least in part on the stored preference, the one or more preferred output devices for output of the second audio content; and sending fourth data for output of the second audio content by the one or more preferred output devices.
2 . The method of claim 1 , further comprising:
performing natural language understanding (NLU) processing on the first data, wherein the determining the first audio content and the one or more output devices is based at least in part on the performing the NLU processing on the first data; and performing the NLU processing on the third data, wherein the determining the second audio content and the determining that the third data omits the specific output device is based at least in part on the performing the NLU processing on the third data.
3 . The method of claim 1 , wherein:
the one or more preferred output devices comprise a group of output devices including a first output device and a second output device; and the sending of the fourth data comprises sending the fourth data to at least one of the first output device or the second output device for synchronized output of the second audio content by the first output device and the second output device.
4 . The method of claim 1 , wherein the fourth data includes a uniform resource locator (URL) associated with a content source of the second audio content.
5 . The method of claim 1 , wherein the fourth data comprises an audio file representing the second audio content.
6 . The method of claim 1 , further comprising, in response to the determining that the third data omits the specific output device:
determining that the first audio content is not being output by the one or more output devices, wherein the determining the stored preference is further based at least in part on the determining that the first audio content is not being output by the one or more output devices.
7 . The method of claim 1 , further comprising, in response to the determining that the third data omits the specific output device:
determining that the third data is not associated with a category of music-related commands, wherein the determining the stored preference is further based at least in part on the determining that the third data is not associated with the category of music-related commands.
8 . A system comprising:
one or more processors; and memory storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to:
determine, based at least in part on first data representing a first utterance, first audio content and one or more output devices;
send second data for output of the first audio content by the one or more output devices;
determine, based at least in part on third data representing a second utterance, second audio content;
determine that the third data omits a specific output device;
determine, based at least in part on determining that the third data omits the specific output device, a stored preference specifying one or more preferred output devices;
select, based at least in part on the stored preference, the one or more preferred output devices for output of the second audio content; and
send fourth data for output of the second audio content by the one or more preferred output devices.
9 . The system of claim 8 , wherein the computer-executable instructions, when executed by the one or more processors, further cause the one or more processors to:
receive the first data from an output device of the one or more output devices; and perform natural language understanding (NLU) processing on the first data, wherein determining the first audio content and the one or more output devices is based at least in part on performing the NLU processing on the first data.
10 . The system of claim 8 , wherein the computer-executable instructions, when executed by the one or more processors, further cause the one or more processors to:
perform natural language understanding (NLU) processing on the third data; and determine an intent to play music based at least in part on performing the NLU processing on the third data, wherein determining the second audio content is based at least in part on determining the intent to play music.
11 . The system of claim 8 , wherein:
the one or more preferred output devices comprise a group of output devices including a first output device and a second output device; and sending the fourth data comprises sending the fourth data to at least one of the first output device or the second output device for synchronized output of the second audio content by the first output device and the second output device.
12 . The system of claim 8 , wherein the fourth data includes a link to an audio file representing the second audio content. 13 . The system of claim 8 , wherein the computer-executable instructions, when executed by the one or more processors, further cause the one or more processors to, in response to determining that the third data omits the specific output device:
determine that the first audio content is not being output by the one or more output devices, wherein determining the stored preference is further based at least in part on determining that the first audio content is not being output by the one or more output devices.
14 . The system of claim 8 , wherein the computer-executable instructions, when executed by the one or more processors, further cause the one or more processors to, in response to determining that the third data omits the specific output device:
determine that the third data is not associated with a category of music-related commands, wherein determining the stored preference is further based at least in part on determining that the third data is not associated with the category of music-related commands.
15 . A method comprising:
determining, based at least in part on first data, first audio content and one or more output devices; sending second data for output of the first audio content by the one or more output devices; determining, based at least in part on third data, second audio content; determining that the third data does not specify any output devices; determining, based at least in part on determining that the third data does not specify any output devices, and based at least in part on a stored preference, a group of output devices for outputting the second audio content, the group of output devices including a first output device and a second output device; and sending fourth data to at least one of the first output device or the second output device for synchronized output of the second audio content by the first output device and the second output device.
16 . The method of claim 15 , wherein:
the one or more output devices comprise a second group of output devices including a third output device and a fourth output device; and the sending of the second data comprises sending the second data to at least one of the third output device or the fourth output device for synchronized output of the first audio content by the third output device and the fourth output device.
17 . The method of claim 15 , wherein the fourth data includes a uniform resource locator (URL) associated with a content streaming source that is usable to stream the second audio content.
18 . The method of claim 15 , wherein the fourth data includes an instruction for an output device of the group of output devices to retrieve, via a local area network (LAN), the second audio content from a content source located in an environment of the group of output devices.
19 . The method of claim 15 , further comprising, in response to the determining that the third data does not specify any output devices:
determining that the third data is not associated with a category of music-related commands, wherein the determining the group of output devices is further based at least in part on the determining that the third data is not associated with the category of music-related commands.
20 . The method of claim 15 , further comprising, in response to the determining that the third data does not specify any output devices:
determining that the first audio content is not being output by the one or more output devices, wherein the determining the group of output devices is further based at least in part on the determining that the first audio content is not being output by the one or more output devices.Join the waitlist — get patent alerts
Track US2021074291A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.