Voice Controlled Media Playback System
Abstract
Disclosed herein are systems and methods for receiving a voice command and determining an appropriate action for the media playback system to execute based on user identification. The systems and methods receive a voice command for a media playback system, and determines whether the voice command was received from a registered user of the media playback system. In response to determining that the voice command was received from a registered user, the systems and methods configure an instruction for the media playback system based on content from the voice command and information in a user profile for the registered user.
Claims
exact text as granted — not AI-modified1 . A system comprising:
a network interface; at least one processor; and at least one non-transitory computer-readable medium comprising program instructions that are executable by the at least one processor such that the system is configured to:
receive, via at least one microphone of at least one network microphone device, first data representing a first voice input comprising a command to play back first audio content;
determine that vocal characteristics of the first voice input correspond to a first profile of a media playback system computing at least one playback device;
select a first streaming audio service from among multiple streaming audio services based on a service preference of the first profile;
cause the at least one playback device to play back first audio content from the selected first streaming audio service according to the first voice input;
receive, via the at least one microphone of the at least one network microphone device, second data representing a second voice input comprising a command to play back second audio content;
determine that vocal characteristics of the first voice input correspond to a second profile of the media playback system;
select a second streaming audio service from among multiple streaming audio services based on a service preference of the second profile; and
cause the at least one playback device to play back second audio content from the selected second streaming audio service according to the second voice input.
2 . The system of claim 1 , wherein the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the system is configured to:
while the at least one playback device is playing back the first audio content, receive, via the at least one microphone of the at least one network microphone device, third data representing a third voice input comprising at least one playback command; determine that vocal characteristics of the third voice input correspond to the second profile of the media playback system; and cause the at least one playback device to carry out the at least one playback command.
3 . The system of claim 2 , wherein the at least one playback command comprises a command to change a playback order of the first audio content, and wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to carry out the at least one playback command comprise program instructions that are executable by the at least one processor such that the system is configured to:
cause the at least one playback device to change a playback order of the first audio content according to the third voice input, wherein the at least one playback device continues playing back the first audio content from the selected first streaming audio service after changing the playback order.
4 . The system of claim 2 , wherein the at least one playback command comprises a command to change a volume setting of the at least one playback device, and wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to carry out the at least one playback command comprise program instructions that are executable by the at least one processor such that the system is configured to:
cause the at least one playback device to change the volume setting of the at least one playback device according to the third voice input, wherein the at least one playback device continues playing back the first audio content from the selected first streaming audio service after changing the volume setting of the at least one playback device.
5 . The system of claim 1 , wherein the vocal characteristics of the first voice input comprise tone characteristics and frequency characteristics, and wherein the program instructions that are executable by the at least one processor such that the system is configured to determine that the vocal characteristics of the first voice input correspond to the first profile of a media playback system comprise program instructions that are executable by the at least one processor such that the system is configured to:
determine that the tone characteristics and frequency characteristics of the first voice input correspond to tone characteristics and frequency characteristics of a voice associated with the first profile.
6 . The system of claim 1 , wherein the system comprises the at least one network microphone device, and wherein the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the system is configured to:
capture, via a plurality of microphones, at least one input sound data stream; monitor the at least one input sound data stream for a wake word of a voice assistant; detect a first instance of the wake word of the voice assistant in a first portion of the at least one input sound data stream; send, to the voice assistant, the first data, wherein the first data represents the first portion of the input sound data stream for processing as the first voice input; detect a second instance of the wake word of the voice assistant in a second portion of the at least one input sound data stream; and send, to the voice assistant, the second data, wherein the second data represents the second portion of the input sound data stream for processing as the second voice input.
7 . The system of claim 6 , wherein the as least one network microphone device comprises a first network microphone device and a second network microphone device, wherein the plurality of microphones comprises at least one first microphone of the first network microphone device and at least one second microphone of the second network microphone device, wherein the program instructions that are executable by the at least one processor such that the system is configured to capture the at least one input sound data stream comprise program instructions that are executable by the at least one processor such that the system is configured to:
capture, via the at least one first microphone, a first input sound data stream of the at least one input sound data stream; and capture, via the at least one second microphone, a second input sound data stream of the at least one input sound data stream.
8 . The system of claim 1 , wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to play back the second audio content from the selected second streaming audio service according to the second voice input comprise program instructions that are executable by the at least one processor such that the system is configured to:
send, via the network interface over at least one network, instructions that cause the at least one playback device to play back the second audio content from the selected second streaming audio service according to the second voice input.
9 . The system of claim 1 , wherein the at least one playback device comprises a first playback device and a second playback device configured in a synchrony group to play back audio in synchrony, and wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to play back the first audio content from the selected first streaming audio service according to the first voice input comprise program instructions that are executable by the at least one processor such that the system is configured to:
send, via the network interface over at least one network, instructions that cause the first playback device and the second playback device to play back the first audio content from the selected first streaming audio service in synchrony according to the second voice input.
10 . The system of claim 1 , wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to play back the first audio content from the selected first streaming audio service according to the first voice input comprise program instructions that are executable by the at least one processor such that the system is configured to:
cause the at least one playback device to stream the first audio content from the selected first streaming audio service using a particular user account of the first streaming audio service associated with the first profile, wherein the program instructions that are executable by the at least one processor such that the system is configured to cause the at least one playback device to play back the second audio content from the selected second streaming audio service according to the second voice input comprise program instructions that are executable by the at least one processor such that the system is configured to:
cause the at least one playback device to stream the second audio content from the selected second streaming audio service using a particular user account of the second streaming audio service associated with the second profile.
11 . At least one computing device comprising:
a network interface; at least one processor; and at least one non-transitory computer-readable medium comprising program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
receive, via at least one microphone of at least one network microphone device, first data representing a first voice input comprising a command to play back first audio content;
determine that vocal characteristics of the first voice input correspond to a first profile of a media playback system comprising at least one playback device;
select a first streaming audio service from among multiple streaming audio services based on a service preference of the first profile;
cause the at least one playback device to play back first audio content from the selected first streaming audio service according to the first voice input;
receive, via the at least one microphone of the at least one network microphone device, second data representing a second voice input comprising a command to play back second audio content;
determine that vocal characteristics of the first voice input correspond to a second profile of the media playback system;
select a second streaming audio service from among multiple streaming audio services based on a service preference of the second profile; and
cause the at least one playback device to play back second audio content from the selected second streaming audio service according to the second voice input.
12 . The at least one computing device of claim 11 , wherein the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
while the at least one playback device is playing back the first audio content, receive, via the at least one microphone of the at least one network microphone device, third data representing a third voice input comprising at least one playback command; determine that vocal characteristics of the third voice input correspond to the second profile of the media playback system; and cause the at least one playback device to carry out the at least one playback command.
13 . The at least one computing device of claim 12 , wherein the at least one playback command comprises a command to change a playback order of the first audio content, and wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to cause the at least one playback device to carry out the at least one playback command comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
cause the at least one playback device to change a playback order of the first audio content according to the third voice input, wherein the at least one playback device continues playing back the first audio content from the selected first streaming audio service after changing the playback order.
14 . The at least one computing device of claim 12 , wherein the at least one playback command comprises a command to change a volume setting of the at least one playback device, and wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to cause the at least one playback device to carry out the at least one playback command comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
cause the at least one playback device to change the volume setting of the at least one playback device according to the third voice input, wherein the at least one playback device continues playing back the first audio content from the selected first streaming audio service after changing the volume setting of the at least one playback device.
15 . The at least one computing device of claim 11 , wherein the vocal characteristics of the first voice input comprise tone characteristics and frequency characteristics, and wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to determine that the vocal characteristics of the first voice input correspond to the first profile of a media playback system comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
determine that the tone characteristics and frequency characteristics of the first voice input correspond to tone characteristics and frequency characteristics of a voice associated with the first profile.
16 . The at least one computing device of claim 11 , wherein the at least one computing device comprises the at least one network microphone device, and wherein the at least one non-transitory computer-readable medium further comprises program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
capture, via a plurality of microphones, at least one input sound data stream; monitor the at least one input sound data stream for a wake word of a voice assistant; detect a first instance of the wake word of the voice assistant in a first portion of the at least one input sound data stream; send, to the voice assistant, the first data, wherein the first data represents the first portion of the input sound data stream for processing as the first voice input; detect a second instance of the wake word of the voice assistant in a second portion of the at least one input sound data stream; and send, to the voice assistant, the second data, wherein the second data represents the second portion of the input sound data stream for processing as the second voice input.
17 . The at least one computing device of claim 16 , wherein the as least one network microphone device comprises a first network microphone device and a second network microphone device, wherein the plurality of microphones comprises at least one first microphone of the first network microphone device and at least one second microphone of the second network microphone device, wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to capture the at least one input sound data stream comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
capture, via the at least one first microphone, a first input sound data stream of the at least one input sound data stream; and capture, via the at least one second microphone, a second input sound data stream of the at least one input sound data stream.
18 . The at least one computing device of claim 11 , wherein the at least one playback device comprises a first playback device and a second playback device configured in a synchrony group to play back audio in synchrony, and wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to cause the at least one playback device to play back the first audio content from the selected first streaming audio service according to the first voice input comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
send, via the network interface over at least one network, instructions that cause the first playback device and the second playback device to play back the first audio content from the selected first streaming audio service in synchrony according to the second voice input.
19 . The at least one computing device of claim 11 , wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to cause the at least one playback device to play back the first audio content from the selected first streaming audio service according to the first voice input comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
cause the at least one playback device to stream the first audio content from the selected first streaming audio service using a particular user account of the first streaming audio service associated with the first profile, wherein the program instructions that are executable by the at least one processor such that the at least one computing device is configured to cause the at least one playback device to play back the second audio content from the selected second streaming audio service according to the second voice input comprise program instructions that are executable by the at least one processor such that the at least one computing device is configured to:
cause the at least one playback device to stream the second audio content from the selected second streaming audio service using a particular user account of the second streaming audio service associated with the second profile.
20 . At least one non-transitory computer-readable medium comprising program instructions that are executable by at least one processor such that at least one computing device is configured to:
receive, via at least one microphone of at least one network microphone device, first data representing a first voice input comprising a command to play back first audio content; determine that vocal characteristics of the first voice input correspond to a first profile of a media playback system comprising at least one playback device; select a first streaming audio service from among multiple streaming audio services based on a service preference of the first profile; cause the at least one playback device to play back first audio content from the selected first streaming audio service according to the first voice input; receive, via the at least one microphone of the at least one network microphone device, second data representing a second voice input comprising a command to play back second audio content; determine that vocal characteristics of the first voice input correspond to a second profile of the media playback system; select a second streaming audio service from among multiple streaming audio services based on a service preference of the second profile; and cause the at least one playback device to play back second audio content from the selected second streaming audio service according to the second voice input.Join the waitlist — get patent alerts
Track US2023297329A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.