Detecting a trigger of a digital assistant
Abstract
Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors, memory, and a plurality of microphones, sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals; processing the plurality of audio signals to obtain a plurality of audio streams; and determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger. The method further includes, in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger, initiating a session of the digital assistant; and in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger, foregoing initiating a session of the digital assistant.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An electronic device, comprising:
one or more processors; a memory; a plurality of microphones; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals;
processing the plurality of audio signals to obtain a plurality of audio streams;
determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger:
in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger: initiating, by the electronic device, a session of a digital assistant;
in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger: foregoing initiating a session of the digital assistant.
2 . The electronic device of claim 1 ,
wherein a first microphone of the plurality of microphones is associated with a first direction, and wherein a second microphone of the plurality of microphones is associated with a second direction different from the first direction.
3 . The electronic device of claim 1 , wherein the plurality of audio streams comprises a plurality of audio beams.
4 . The electronic device of claim 1 , wherein processing the plurality of audio signals to obtain a plurality of audio streams comprises processing an audio signal of the plurality of audio signals using source separation.
5 . The electronic device of claim 1 , wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining whether each of the plurality of audio streams includes the spoken trigger.
6 . The electronic device of claim 1 , wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining whether a combination of two more audio streams of the plurality of audio streams includes the spoken trigger.
7 . The electronic device of claim 1 , the one or more programs further including instructions for:
obtaining one or more trigger scores corresponding to the plurality of audio streams; wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining whether any of the plurality of audio signals corresponds to the spoken trigger based on the one or more trigger scores.
8 . The electronic device of claim 1 , wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining whether any of the plurality of audio signals corresponds to the spoken trigger based on acoustic information associated with a user of the electronic device.
9 . The electronic device of claim 1 , the one or more programs further including instructions for:
obtaining a plurality of words based on the plurality of audio streams; wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining whether any of the plurality of audio signals corresponds to the spoken trigger based on information corresponding to the plurality of words.
10 . The electronic device of claim 9 , the one or more programs further including instructions for:
obtaining one or more parse results based on the plurality of words; wherein the information corresponding to the plurality of words includes the one or more parse results.
11 . The electronic device of claim 9 , the one or more programs further including instructions for:
obtaining one or more representations of user intent based on the plurality of words; wherein the information corresponding to the plurality of words includes the one or more representations of user intent.
12 . The electronic device of claim 9 , wherein the information corresponding to the plurality of words is indicative of a direction.
13 . The electronic device of claim 9 , wherein the information corresponding to the plurality of words is indicative of a speaker.
14 . The electronic device of claim 1 , wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
identifying, from the plurality of audio streams, a set of candidate audio streams; providing, to a remote device, one or more candidate audio streams from the set of candidate audio streams; and obtaining validation information from the remote device.
15 . The electronic device of claim 14 , the one or more programs further including instructions for:
selecting the one or more candidate audio streams from the set of candidate audio streams based on respective trigger scores associated with the one or more candidate audio streams.
16 . The electronic device of claim 15 , the one or more programs further including instructions for:
providing each of the plurality of audio streams to a neural network to obtain a respective trigger score.
17 . The electronic device of claim 14 , the one or more programs further including instructions for:
selecting the one or more candidate audio streams from the set of candidate audio streams based on respective entropy information associated with the one or more candidate audio streams.
18 . The electronic device of claim 14 , the one or more programs further including instructions for:
determining that a first candidate audio stream corresponds to a spoken trigger detected at a first time; determining, at a second time, that a second candidate audio stream corresponds to a spoken trigger detected at a second time; and selecting the one or more candidate audio streams from the set of candidate audio streams based on the first time and the second time.
19 . The electronic device of claim 1 , the one or more programs further including instructions for:
in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger:
identifying a first segment of an audio stream of the plurality of audio streams;
identifying a second segment of the audio stream;
determining whether the first segment and the second segment correspond to a same user.
20 . The electronic device of claim 19 , the one or more programs further including instructions for:
in accordance with a determination that the first segment and the second segment correspond to the same user:
determining that the user is a user of the electronic device; and
obtaining a representation of user intent based on the first segment and the second segment.
21 . The electronic device of claim 19 , wherein determining whether the first segment and the second segment correspond to a same user comprises:
comparing acoustic information associated with the first segment with acoustic information associated with the second segment.
22 . The electronic device of claim 19 , wherein determining whether the first segment and the second segment correspond to a same user comprises:
identifying a first entropy associated with the first segment; identifying a second entropy associated with the second segment; and comparing the first entropy with the second entropy.
23 . The electronic device of claim 19 , wherein determining whether the first segment and the second segment correspond to a same user comprises:
identifying a parse result based on the first segment and the second segment.
24 . The electronic device of claim 1 , wherein the electronic device is a first electronic device, the one or more programs further including instructions for:
receiving, from a second electronic device, information corresponding to an audio signal detected at the second electronic device; wherein determining whether any of the plurality of audio signals corresponds to the spoken trigger comprises:
determining, based on the information received from the second electronic device, whether the audio signal detected at the second electronic device corresponds to the spoken trigger.
25 . The electronic device of claim 24 , wherein the information includes location information of one or more microphones of the second electronic device.
26 . The electronic device of claim 24 , wherein the information includes directional information of the audio signal detected at the second electronic device.
27 . The electronic device of claim 24 , wherein the information includes a device type associated with the second electronic device.
28 . The electronic device of claim 24 , wherein the second electronic device is associated with a different device type than the first electronic device.
29 . The electronic device of claim 1 , wherein initiating a session of the digital assistant comprises:
providing, by the digital assistant, an audio output.
30 . The electronic device of claim 29 , wherein each of the plurality of audio streams is associated with directional information, and wherein providing an audio output comprises:
providing, by the digital assistant, the audio output based on the directional information associated with the plurality of audio streams.
31 . The electronic device of claim 1 , wherein the each of the plurality of audio streams is associated with directional information, and wherein determining whether any of the plurality of audio signals corresponds to a spoken trigger comprises:
determining, based on the plurality of audio streams and associated directional information, whether any of the plurality of audio signals corresponds to a spoken trigger.
32 . The electronic device of claim 1 ,
wherein the plurality of audio signals is the first plurality of audio signals, wherein determining whether any of the plurality of audio signals corresponds to a spoken trigger comprises: detecting the spoken trigger from one or more audio streams of the plurality of audio streams, the one or more programs further including instructions for:
selecting, based on the one or more audio streams, a set of microphones of the plurality of microphones; and
sampling, using the first set of microphones, a second plurality of audio signals.
33 . The electronic device of claim 1 , wherein the electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or a combination thereof.
34 . A method for operating a digital assistant, comprising:
at an electronic device with one or more processors, memory, and a plurality of microphones:
sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals;
processing the plurality of audio signals to obtain a plurality of audio streams;
determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger:
in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger: initiating, by the electronic device, a session of the digital assistant;
in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger: foregoing initiating a session of the digital assistant.
35 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device with a plurality of microphones, cause the electronic device to:
sample, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals; process the plurality of audio signals to obtain a plurality of audio streams; determine, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger: in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger: initiate, by the electronic device, a session of a digital assistant; in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger: forego initiating a session of the digital assistant.Join the waitlist — get patent alerts
Track US2018336892A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.