Intercom-style communication using multiple computing devices
Abstract
Techniques are described related to improved intercom-style communication using a plurality of computing devices distributed about an environment. In various implementations, voice input may be received, e.g., at a microphone of a first computing device of multiple computing devices, from a first user. The voice input may be analyzed and, based on the analyzing, it may be determined that the first user intends to convey a message to a second user. A location of the second user relative to the multiple computing devices may be determined, so that, based on the location of the second user, a second computing device may be selected from the multiple computing devices that is capable of providing audio or visual output that is perceptible to the second user. The second computing device may then be operated to provide audio or visual output that conveys the message to the second user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
accessing a trained machine learning model, wherein the machine learning model is trained, using a corpus of labeled voice inputs, to predict whether voice inputs are indicative of background conversation that should be ignored, or are indicative of a user intent to convey a message to one or more other users; receiving, at a microphone of a first computing device of a plurality of computing devices, from a first user, voice input; analyzing the voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate output, wherein the output indicates that the first user intends to convey a message to the one or more other users; determining, based on the analyzing, that the first user intends to convey the message to the one or more other users; and causing one or more other computing devices of the plurality of computing devices to provide audio or visual output that conveys the message to the one or more other users.
2 . The method of claim 1 , further comprising:
receiving, at the microphone of the first computing device, an additional voice input; analyzing the additional voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate additional output, wherein the additional output indicates that the additional voice input is indicative of background noise that should be ignored; ignoring the additional voice input in response to the additional output indicating that the additional voice input is indicative of background noise that should be ignored.
3 . The method of claim 1 , wherein the machine learning model is trained using a corpus of labeled voice inputs, and wherein labels applied to the voice inputs include:
a first label indicative of a user intent to convey a message to one or more other users; and a second label indicative of background conversation between multiple users.
4 . The method of claim 3 , wherein the labels applied to the voice inputs further include a third label indicative of a user intent to engage in a human-to-computer dialog with an automated assistant.
5 . The method of claim 1 , further comprising:
determining a location of a second user of the one or more users relative to the plurality of computing devices; and selecting, from the plurality of computing devices, based on the location of the second user, a second computing device that is capable of providing audio or visual output that is perceptible to the second user; wherein the causing includes causing the second computing device provide the audio or visual output that conveys the message to the second user.
6 . The method of claim 1 , wherein the causing comprising broadcasting the message to the one or more other users using all of the plurality of computing devices.
7 . The method of claim 1 , wherein the analyzing includes performing speech-to-text processing on an audio recording of the voice input to generate, as the data indicative of the audio recording, textual input, wherein the textual input is applied as input across the trained machine learning.
8 . A system comprising one or more processors and memory operably coupled with the one or more processors, wherein the memory stores instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to perform the following operations:
accessing a trained machine learning model, wherein the machine learning model is trained, using a corpus of labeled voice inputs, to predict whether voice inputs are indicative of background conversation that should be ignored, or are indicative of a user intent to convey a message to one or more other users; receiving, at a microphone of a first computing device of a plurality of computing devices, from a first user, voice input; analyzing the voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate output, wherein the output indicates that the first user intends to convey a message to the one or more other users; determining, based on the analyzing, that the first user intends to convey the message to the one or more other users; and causing one or more other computing devices of the plurality of computing devices to provide audio or visual output that conveys the message to the one or more other users.
9 . The system of claim 8 , further comprising:
receiving, at the microphone of the first computing device, an additional voice input; analyzing the additional voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate additional output, wherein the additional output indicates that the additional voice input is indicative of background noise that should be ignored; ignoring the additional voice input in response to the additional output indicating that the additional voice input is indicative of background noise that should be ignored.
10 . The system of claim 8 , wherein the machine learning model is trained using a corpus of labeled voice inputs, and wherein labels applied to the voice inputs include:
a first label indicative of a user intent to convey a message to one or more other users; and a second label indicative of background conversation between multiple users.
11 . The system of claim 10 , wherein the labels applied to the voice inputs further include a third label indicative of a user intent to engage in a human-to-computer dialog with an automated assistant.
12 . The system of claim 8 , further comprising:
determining a location of a second user of the one or more users relative to the plurality of computing devices; and selecting, from the plurality of computing devices, based on the location of the second user, a second computing device that is capable of providing audio or visual output that is perceptible to the second user; wherein the causing includes causing the second computing device provide the audio or visual output that conveys the message to the second user.
13 . The system of claim 8 , wherein the causing comprising broadcasting the message to the one or more other users using all of the plurality of computing devices.
14 . The system of claim 8 , wherein the analyzing includes performing speech-to-text processing on an audio recording of the voice input to generate, as the data indicative of the audio recording, textual input, wherein the textual input is applied as input across the trained machine learning.
15 . At least one non-transitory computer-readable medium comprising instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to perform the following operations:
accessing a trained machine learning model, wherein the machine learning model is trained, using a corpus of labeled voice inputs, to predict whether voice inputs are indicative of background conversation that should be ignored, or are indicative of a user intent to convey a message to one or more other users; receiving, at a microphone of a first computing device of a plurality of computing devices, from a first user, voice input; analyzing the voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate output, wherein the output indicates that the first user intends to convey a message to the one or more other users; determining, based on the analyzing, that the first user intends to convey the message to the one or more other users; and causing one or more other computing devices of the plurality of computing devices to provide audio or visual output that conveys the message to the one or more other users.
16 . The at least one non-transitory computer-readable medium of claim 15 , further comprising instructions for:
receiving, at the microphone of the first computing device, an additional voice input; analyzing the additional voice input, wherein the analyzing includes applying data indicative of an audio recording of the voice input as input across the trained machine learning model to generate additional output, wherein the additional output indicates that the additional voice input is indicative of background noise that should be ignored; ignoring the additional voice input in response to the additional output indicating that the additional voice input is indicative of background noise that should be ignored.
17 . The at least one non-transitory computer-readable medium of claim 15 , wherein the machine learning model is trained using a corpus of labeled voice inputs, and wherein labels applied to the voice inputs include:
a first label indicative of a user intent to convey a message to one or more other users; and a second label indicative of background conversation between multiple users.
18 . The at least one non-transitory computer-readable medium of claim 17 , wherein the labels applied to the voice inputs further include a third label indicative of a user intent to engage in a human-to-computer dialog with an automated assistant.
19 . The at least one non-transitory computer-readable medium of claim 15 , further comprising:
determining a location of a second user of the one or more users relative to the plurality of computing devices; and selecting, from the plurality of computing devices, based on the location of the second user, a second computing device that is capable of providing audio or visual output that is perceptible to the second user; wherein the causing includes causing the second computing device provide the audio or visual output that conveys the message to the second user.
20 . The at least one non-transitory computer-readable medium of claim 15 , wherein the causing comprising broadcasting the message to the one or more other users using all of the plurality of computing devices.Join the waitlist — get patent alerts
Track US2019079724A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.