Securely Executing Voice Actions Using Contextual Signals
Abstract
In some implementations, (i) audio data representing a voice command spoken by a speaker and (ii) a speaker identification result indicating that the voice command was spoken by the speaker are obtained. A voice action is selected based at least on a transcription of the audio data. A service provider corresponding to the selected voice action is selected from among a plurality of different service providers. One or more input data types that the selected service provider uses to perform authentication for the selected voice action are identified. A request to perform the selected voice action and (i) one or more values that correspond to the identified one or more input data types are provided to the service provider.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:
receiving audio data representing a voice command spoken by a speaker and captured by a first computing device; obtaining, based on the audio data representing the voice command, a speaker identification result indicating that the voice command was spoken by the speaker; determining, based on the speaker identification result indicating that the voice command was spoken by the speaker, a device identifier that corresponds to a second computing device; selecting a voice action based on a transcription of the audio data; determining, using contextual data from the second computing device, that the second computing device coincides with a location designated by the speaker; and providing, to a service provider, a request to perform the selected voice action.
2 . The method of claim 1 , wherein obtaining the speaker identification result indicating that the voice command was spoken by the speaker is based on a likelihood that the audio data representing the voice command matches a stored voice print associated with the speaker.
3 . The method of claim 1 , wherein obtaining the speaker identification result indicating that the voice command was spoken by the speaker comprises:
obtaining a plurality of speaker identification results each indicating a corresponding likelihood that the audio data representing the voice command matches a corresponding one of a plurality of stored voice prints associated with different speakers; and selecting, from among the plurality of speaker identification results, the speaker identification result having the highest corresponding likelihood as the speaker identification result indicating that the voice command was spoken by the speaker.
4 . The method of claim 1 , wherein the operations further comprise, after selecting the voice action based on the transcription of the audio data, selecting, from a plurality of different service providers, the service provider that can perform the selected voice action.
5 . The method of claim 4 , wherein selecting the service provider that can perform the selected voice action comprises:
obtaining a mapping of voice actions to the plurality of different service providers, each voice action the mapping describing a service provider that can perform the voice action; determining that the mapping of voice actions indicates that the service provider can perform the selected voice action; and in response to determining that the mapping of voice actions indicates that the service provider can perform the selected voice action, selecting the service provider.
6 . The method of claim 1 , wherein selecting the voice action based on the transcription of the audio data comprises:
obtaining a set of voice actions, wherein each voice action in the set of voice actions identifies one or more terms that correspond to that voice action; determining that one or more terms in the transcription match the one or more terms that correspond to the voice action; and in response to determining that the one or more terms in the transcription match the one or more terms that correspond to the voice action, selecting the voice action from among the set of voice actions.
7 . The method of claim 1 , further comprising generating the transcription of the audio data using an automated speech recognizer.
8 . The method of claim 1 , further comprising receiving, from the service provider, an indication that the service provider performed the selected voice action.
9 . The method of claim 8 , wherein the first computing device is configured to output the indication that the service provider performed the selected voice action as synthesized speech.
10 . The method of claim 1 , wherein the operations further comprise, after providing the request to perform the selected voice action to the service provider, providing, to the first computing device, an authorization request requesting the speaker to provide an explicit authorization code that the service provider needs to perform the selected voice action.
11 . A system comprising:
data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving audio data representing a voice command spoken by a speaker and captured by a first computing device;
obtaining, based on the audio data representing the voice command, a speaker identification result indicating that the voice command was spoken by the speaker;
determining, based on the speaker identification result indicating that the voice command was spoken by the speaker, a device identifier that corresponds to a second computing device;
selecting a voice action based on a transcription of the audio data;
determining, using contextual data from the second computing device, that the second computing device coincides with a location designated by the speaker; and
providing, to a service provider, a request to perform the selected voice action.
12 . The system of claim 11 , wherein obtaining the speaker identification result indicating that the voice command was spoken by the speaker is based on a likelihood that the audio data representing the voice command matches a stored voice print associated with the speaker.
13 . The system of claim 11 , wherein obtaining the speaker identification result indicating that the voice command was spoken by the speaker comprises:
obtaining a plurality of speaker identification results each indicating a corresponding likelihood that the audio data representing the voice command matches a corresponding one of a plurality of stored voice prints associated with different speakers; and selecting, from among the plurality of speaker identification results, the speaker identification result having the highest corresponding likelihood as the speaker identification result indicating that the voice command was spoken by the speaker.
14 . The system of claim 11 , wherein the operations further comprise, after selecting the voice action based on the transcription of the audio data, selecting, from a plurality of different service providers, the service provider that can perform the selected voice action.
15 . The system of claim 14 , wherein selecting the service provider that can perform the selected voice action comprises:
obtaining a mapping of voice actions to the plurality of different service providers, each voice action the mapping describing a service provider that can perform the voice action; determining that the mapping of voice actions indicates that the service provider can perform the selected voice action; and in response to determining that the mapping of voice actions indicates that the service provider can perform the selected voice action, selecting the service provider.
16 . The system of claim 11 , wherein selecting the voice action based on the transcription of the audio data comprises:
obtaining a set of voice actions, wherein each voice action in the set of voice actions identifies one or more terms that correspond to that voice action; determining that one or more terms in the transcription match the one or more terms that correspond to the voice action; and in response to determining that the one or more terms in the transcription match the one or more terms that correspond to the voice action, selecting the voice action from among the set of voice actions.
17 . The system of claim 11 , further comprising generating the transcription of the audio data using an automated speech recognizer.
18 . The system of claim 11 , further comprising receiving, from the service provider, an indication that the service provider performed the selected voice action.
19 . The system of claim 18 , wherein the first computing device is configured to output the indication that the service provider performed the selected voice action as synthesized speech.
20 . The system of claim 11 , wherein the operations further comprise, after providing the request to perform the selected voice action to the service provider, providing, to the first computing device, an authorization request requesting the speaker to provide an explicit authorization code that the service provider needs to perform the selected voice action.Join the waitlist — get patent alerts
Track US2023269586A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.