Remote system processing based on a previously identified user
Abstract
Techniques for implementing a “sticky” user ID are described. A system receives first input audio data and determines first speech processing results therefrom. The system also determines a first user ID of a user that spoke an utterance represented in the first input audio data and associates the first user ID with a device, which originated the first input audio data, for a predetermined length of time. The system determines first output data responsive to the first speech processing data and causes the device to present first output content corresponding thereto. The system then receives second input audio data and determines second speech processing results therefrom. The system also determines a time of receipt of the second input audio data is within the predetermined length of time. Based at least in part thereon, the system determined second output data responsive to the second speech processing data using the first user ID. The system then causes the device to present second output content corresponding to the second output data.
Claims
exact text as granted — not AI-modified1 .- 20 . (canceled)
21 . A computer-implemented method, comprising:
receiving, from a device, first audio data corresponding to an utterance; performing user recognition processing on the first audio data to determine the utterance was spoken by a user corresponding to a first user type; performing speech processing on the first audio data to determine a request to perform an action; determining, based at least in part on the first user type, output data responsive to the request; and causing presentation of an output using the output data.
22 . The computer-implemented method of claim 21 , further comprising:
determining control rules corresponding to the first user type, wherein determining the output data is based at least in part on the control rules.
23 . The computer-implemented method of claim 21 , wherein performing the speech processing comprises performing automatic speech recognition (ASR) using the first audio data to determine ASR data, and wherein the method further comprises:
based at least in part on the first user type, sending the ASR data to a first application corresponding to the first user type.
24 . The computer-implemented method of claim 21 , further comprising:
determining the request corresponds to a first application; sending, to a component associated with the first application, an indication of the first user type; and receiving, from the component, the output data, wherein the output data was determined based at least in part on the first user type.
25 . The computer-implemented method of claim 21 , further comprising:
sending, to a first application, an indication of the request; sending, to a second application, an indication of the request; receiving first response data from the first application, the first response data corresponding to the request; receiving second response data from the second application, the second response data corresponding to the request; and based at least in part on the first user type, including the first response data rather than the second response data in the output data.
26 . The computer-implemented method of claim 21 , further comprising:
determining the action is not performable in response to the request corresponding to the first user type; including in the output data an indication corresponding to further authorization; receiving second audio data corresponding to a second utterance; performing user recognition processing on the second audio data to determine the utterance was spoken by a second user corresponding to a second user type; determining the action is performable in response to the second utterance corresponding to the second user type; and causing execution of the action.
27 . The computer-implemented method of claim 26 , wherein the second utterance does not include a wakeword.
28 . The computer-implemented method of claim 21 , wherein:
performing the speech processing comprises determining first NLU data corresponding to a first intent and second NLU data corresponding to a second intent; and determining the output data comprises:
determining the first NLU data corresponds to the first user type, and
causing processing to be performed using the first NLU data to determine the output data.
29 . The computer-implemented method of claim 21 , further comprising:
after performing the user recognition processing, associating the first user type with the device.
30 . The computer-implemented method of claim 21 , wherein the first user type corresponds to a guest type corresponding to the device.
31 . A system comprising:
at least one processor; and at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive, from a device, first audio data corresponding to an utterance;
perform user recognition processing on the first audio data to determine the utterance was spoken by a user corresponding to a first user type;
perform speech processing on the first audio data to determine a request to perform an action;
determine, based at least in part on the first user type, output data responsive to the request; and
cause presentation of an output using the output data.
32 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine control rules corresponding to the first user type, wherein determination of the output data is based at least in part on the control rules.
33 . The system of claim 31 , wherein the instructions that cause the system to perform the speech processing comprise instructions that, when executed by the at least one processor, cause the system to perform automatic speech recognition (ASR) using the first audio data to determine ASR data, and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
based at least in part on the first user type, send the ASR data to a first application corresponding to the first user type.
34 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the request corresponds to a first application; send, to a component associated with the first application, an indication of the first user type; and receive, from the component, the output data, wherein the output data was determined based at least in part on the first user type.
35 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
send, to a first application, an indication of the request; send, to a second application, an indication of the request; receive first response data from the first application, the first response data corresponding to the request; receive second response data from the second application, the second response data corresponding to the request; and based at least in part on the first user type, include the first response data rather than the second response data in the output data.
36 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine the action is not performable in response to the request corresponding to the first user type; include in the output data an indication corresponding to further authorization; receive second audio data corresponding to a second utterance; perform user recognition processing on the second audio data to determine the utterance was spoken by a second user corresponding to a second user type; determine the action is performable in response to the second utterance corresponding to the second user type; and cause execution of the action.
37 . The system of claim 36 , wherein the second utterance does not include a wakeword.
38 . The system of claim 31 , wherein:
the instructions that cause the system to perform the speech processing comprise instructions that, when executed by the at least one processor, cause the system to determine first NLU data corresponding to a first intent and second NLU data corresponding to a second intent; and the instructions that cause the system to determine the output data comprise instructions that, when executed by the at least one processor, cause the system to:
determine the first NLU data corresponds to the first user type, and
cause processing to be performed using the first NLU data to determine the output data.
39 . The system of claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
after performing the user recognition processing, associating the first user type with the device.
40 . The system of claim 31 , wherein the first user type corresponds to a guest type corresponding to the device.Join the waitlist — get patent alerts
Track US2023388382A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.