Detecting a trigger of a digital assistant
Abstract
Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors, memory, and a plurality of microphones, sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals; processing the plurality of audio signals to obtain a plurality of audio streams; and determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger. The method further includes, in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger, initiating a session of the digital assistant; and in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger, foregoing initiating a session of the digital assistant.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . One or more non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of one or more electronic devices, cause the one or more electronic devices to:
sample, using a first microphone of a first electronic device, a first audio signal; sample, using a second microphone of a second electronic device different from the first electronic device, a second audio signal; determine, at a third electronic device different from the first electronic device and the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger; obtain directional information associated with an audio source based on the first audio signal and the second audio signal;
in accordance with a determination that the first audio signal or the second audio signal correspond to a spoken trigger:
initiate, by a fourth electronic device, a session of the digital assistant, wherein initiating the session of the digital assistant comprises providing, by the digital assistant, an audio output based on the directional information;
in accordance with a determination that the first audio signal and the second audio signal do not correspond to the spoken trigger:
forego, by the fourth electronic device, initiating a session of the digital assistant.
2 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein the fourth electronic device is the first electronic device.
3 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein the third electronic device is a remote server device.
4 . The one or more non-transitory computer-readable storage medium of claim 1 , the one or more programs further comprising instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
obtain, by the third electronic device, information corresponding to the first audio signal and the second audio signal, wherein the information includes location information of the first microphone.
5 . The one or more non-transitory computer-readable storage medium of claim 1 , the one or more programs further comprising instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
obtain, by the third electronic device, information corresponding to the first audio signal and the second audio signal, wherein the information includes directional information of the first audio signal.
6 . The one or more non-transitory computer-readable storage medium of claim 1 , the one or more programs further comprising instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
obtain, by the third electronic device, information corresponding to the first audio signal and the second audio signal, wherein the information includes a device type associated with the first electronic device.
7 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein the second electronic device is a device of a first type and the third electronic device is a device of a second type different from the first type.
8 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to a spoken trigger comprises: determining whether a combination of the first audio signal and the second audio signal corresponds to a spoken trigger.
9 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein the first electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or any combination thereof.
10 . The one or more non-transitory computer-readable storage medium of claim 1 , wherein the second electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or any combination thereof.
11 . A system, comprising:
one or more processors of one or more electronic devices; one or more memories of the one or more electronic devices; and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:
sampling, using a first microphone of a first electronic device, a first audio signal;
sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal;
determining, at a third electronic device different from the first electronic device and the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger;
obtaining directional information associated with an audio source based on the first audio signal and the second audio signal;
in accordance with a determination that the first audio signal or the second audio signal correspond to a spoken trigger:
initiating, by a fourth electronic device, a session of the digital assistant, wherein initiating the session of the digital assistant comprises providing, by the digital assistant, an audio output based on the directional information;
in accordance with a determination that the first audio signal and the second audio signal do not correspond to the spoken trigger:
foregoing, by the fourth electronic device, initiating a session of the digital assistant
12 . A method for operating a digital assistant, the method comprising:
sampling, using a first microphone of a first electronic device, a first audio signal; sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal; determining, at a third electronic device different from the first electronic device and the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger; obtaining directional information associated with an audio source based on the first audio signal and the second audio signal; in accordance with a determination that the first audio signal or the second audio signal correspond to a spoken trigger:
initiating, by a fourth electronic device, a session of the digital assistant, wherein initiating the session of the digital assistant comprises providing, by the digital assistant, an audio output based on the directional information;
in accordance with a determination that the first audio signal and the second audio signal do not correspond to the spoken trigger:
foregoing, by the fourth electronic device, initiating a session of the digital assistant.Join the waitlist — get patent alerts
Track US2019074009A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.