Interactive voice controlled entertainment
Abstract
Methods and systems for receiving shouted-out user responses to broadcast entertainment content, and for determining the responsiveness of those responses in relation to the broadcast content. In particular, entertainment broadcasts can be accompanied by mark-up data that represents various events within a given broadcast, which can be compared to the shouted-out responses to determine their accuracy. For example, if a game show was broadcast and an individual started shouting out answers during the broadcast, embodiments disclosed herein could utilize a voice-controlled electronic device that captures the shouted-out answers and passes them on to a language processing system that determines whether they are correct by comparing the answers to the mark-up data. The voice-controlled electronic device can also “listen” to background sounds to capture the broadcast of the entertainment content, and send that content to the language processing system, which can use that captured data to synchronize the actual broadcast with the analysis of the shouted-out answers to provide individuals with an immersive entertainment experience.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, comprising:
operating a device in a first mode corresponding to sending audio data for language processing following detection of a wakeword; receiving first audio data representing a first utterance; processing the first audio data to determine a representation of the wakeword; in response to determining the representation of the wakeword, causing language processing to be performed using at least a portion of the first audio data; receiving an indication to operate the device in a second mode corresponding sending audio data for language processing without detection of a wakeword; in response to receiving the indication, operating the device in the second mode; receiving second audio data representing a second utterance; and in response to operating the device in the second mode, causing language processing to be performed using at least a portion of the second audio data regardless of whether the second audio data includes a representation of the wakeword.
22 . The computer-implemented method of claim 21 , further comprising:
determining the first utterance corresponds to a session involving a skill; determining the session is to involve further utterances; and in response to the session involving further utterances, generating the indication.
23 . The computer-implemented method of claim 21 , further comprising:
determining the second utterance is directed to the device.
24 . The computer-implemented method of claim 21 , further comprising:
determining a skill operating with respect to the device; and operating the device in the second mode for utterances corresponding to the skill.
25 . The computer-implemented method of claim 24 , further comprising:
receiving third audio data representing a third utterance not related to the skill; and prior to causing language processing to be performed using at least a portion of the third audio data, determining the third audio data includes a representation of the wakeword.
26 . The computer-implemented method of claim 24 , further comprising:
prior to causing the language processing to be performed using at least the portion of the second audio data, determining the second audio data corresponds to the skill.
27 . The computer-implemented method of claim 24 , further comprising:
receiving third audio data representing a third utterance not related to the skill; and discarding the third audio data.
28 . The computer-implemented method of claim 27 , further comprising, by the device:
determining that the third utterance is not related to the skill.
29 . The computer-implemented method of claim 24 , further comprising:
determining the skill has ceased operation with respect to the device; and discontinuing operation of the device in the second mode.
30 . The computer-implemented method of claim 21 , wherein the first utterance was spoken by a first user and the method further comprises:
receiving third audio data representing a third utterance spoken by a second user; and prior to causing language processing to be performed using at least a portion of the third audio data, determining the third audio data includes a representation of the wakeword.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
operate a device in a first mode corresponding to sending audio data for language processing following detection of a wakeword;
receive first audio data representing a first utterance;
process the first audio data to determine a representation of the wakeword;
in response to determining the representation of the wakeword, cause language processing to be performed using at least a portion of the first audio data;
receive an indication to operate the device in a second mode corresponding sending audio data for language processing without detection of a wakeword;
in response to receiving the indication, operating the device in the second mode;
receive second audio data representing a second utterance; and
in response to operating the device in the second mode, cause language processing to be performed using at least a portion of the second audio data regardless of whether the second audio data includes a representation of the wakeword.
32 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the first utterance corresponds to a session involving a skill; determine the session is to involve further utterances; and in response to the session involving further utterances, generate the indication.
33 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the second utterance is directed to the device.
34 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine a skill operating with respect to the device; and operate the device in the second mode for utterances corresponding to the skill.
35 . The system of claim 34 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third audio data representing a third utterance not related to the skill; and prior to causing language processing to be performed using at least a portion of the third audio data, determine the third audio data includes a representation of the wakeword.
36 . The system of claim 34 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
prior to causing the language processing to be performed using at least the portion of the second audio data, determine the second audio data corresponds to the skill.
37 . The system of claim 34 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third audio data representing a third utterance not related to the skill; and discard the third audio data.
38 . The system of claim 37 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
cause the device to determine that the third utterance is not related to the skill.
39 . The system of claim 34 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the skill has ceased operation with respect to the device; and discontinue operation of the device in the second mode.
40 . The system of claim 31 , the first utterance was spoken by a first user and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third audio data representing a third utterance spoken by a second user; and prior to causing language processing to be performed using at least a portion of the third audio data, determine the third audio data includes a representation of the wakeword.Join the waitlist — get patent alerts
Track US2021280185A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.