US12417766B2ActiveUtilityA1
Voice user interface using non-linguistic input
Est. expirySep 30, 2040(~14.2 yrs left)· nominal 20-yr term from priority
Inventors:Colby Nelson Leider
G10L 2015/227G10L 25/63G10L 15/1807G10L 15/005G10L 15/22G06F 3/167
62
PatentIndex Score
0
Cited by
275
References
21
Claims
Abstract
A voice user interface (VUI) and methods for operating the VUI are disclosed. In some embodiments, the VUI configured to receive and process linguistic and non-linguistic inputs. For example, the VUI receives an audio signal, and the VUI determines whether the audio input comprises a linguistic and/or a non-linguistic input. In accordance with a determination that the audio signal comprises a non-linguistic input, the VUI causes a system to perform an action associated with the non-linguistic input.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A system comprising:
a plurality of microphones of a wearable head device; and
one or more processors configured to perform a method comprising:
receiving, using the plurality of microphones, an audio signal;
determining whether the audio signal is associated with a user of the wearable head device, wherein the determining whether the audio signal is associated with the user comprises:
determining respective distances between a source of the audio signal and each of the plurality of microphones, and
determining, based on the respective distances, a location of the source of the audio signal;
in accordance with a determination that the audio signal is associated with the user of the wearable head device:
determining whether the audio signal comprises a non-linguistic input; and
in accordance with a determination that the audio signal comprises the non-linguistic input, performing an action associated with the non-linguistic input; and
in accordance with a determination that the audio signal is not associated with the user of the wearable head device, forgoing determining whether the audio signal comprises the non-linguistic input.
2. The system of claim 1 , wherein:
the non-linguistic input comprises a paralinguistic input, and
in accordance with a determination that the audio signal comprises the paralinguistic input, the action comprises a first action associated with the paralinguistic input.
3. The system of claim 1 , wherein:
the non-linguistic input comprises a prosodic input, and
in accordance with a determination that the audio signal comprises the prosodic input, the action comprises a second action associated with the prosodic input.
4. The system of claim 3 , wherein the second action is a modification of an action associated with a linguistic input.
5. The system of claim 3 , wherein:
the prosodic input is indicative of an emotion associated with the user of the wearable head device, and
the action is further associated with the emotion.
6. The system of claim 1 , wherein the method further comprises:
determining whether the audio signal comprises a linguistic input; and
in accordance with a determination that the audio signal comprises the linguistic input, performing a third action associated with the linguistic input.
7. The system of claim 6 , wherein the action comprises a modification of the third action based on the non-linguistic input.
8. The system of claim 1 , wherein the method further comprises: in accordance with a determination that the audio signal comprises the non-linguistic input, receiving information associated with the audio signal from a convolutional neural network, wherein the action is performed based on the information.
9. The system of claim 1 , wherein the method further comprises:
classifying a feature of the non-linguistic input, wherein the action is performed based on the classified feature.
10. The system of claim 1 , wherein the method further comprises:
associating the action with the non-linguistic input.
11. The system of claim 1 , wherein the action comprises one of texting, performing an intent, and inserting an emoji.
12. The system of claim 1 , further comprising: a sensor different from the microphone, wherein the method further comprises receiving, from the sensor, information associated with an environment of the system, wherein the action is further associated with the information received from the sensor.
13. The system of claim 1 , wherein:
the system comprises a mixed reality system in a mixed reality environment,
the mixed reality system comprises the wearable head device, and
the action is further associated with the mixed reality environment.
14. The system of claim 1 , wherein: the method further comprises determining a position of the system, wherein the action is further associated with the position of the system.
15. The system of claim 1 , wherein:
in accordance with a determination that the system is associated with a first user, the action comprises a first action associated with the first user; and
in accordance with a determination that the system is associated with a second user, different from the first user, the action comprises a second action associated with the second user, different from the first action.
16. The system of claim 1 , wherein:
the audio signal comprises a frequency-domain feature and a time-domain feature, and
the determination of whether the audio comprises the non-linguistic input is based on the frequency-domain feature and the time-domain feature.
17. The system of claim 1 , wherein the method further comprises:
receiving information from a feature database, wherein the determination of whether the audio comprises the non-linguistic input is further based on the information.
18. The system of claim 1 , wherein determining whether the audio signal comprises the non-linguistic input comprises using a first processor, and the method further comprises: in accordance with a determination that the audio signal comprises the non-linguistic input, waking up a second processor to perform the action.
19. The system of claim 1 , wherein:
the plurality of microphones comprises a first microphone and a second microphone,
the first microphone is at a first location of the wearable head device, the first microphone configured to rest at a first distance from a user's mouth, and
the second microphone is at a second location of the wearable head device, the second microphone configured to rest at a second distance from the user's mouth, the second distance unequal to the first distance.
20. A method comprising:
receiving, using a plurality of microphones of a wearable head device, an audio signal;
determining whether the audio signal is associated with a user of the wearable head device, wherein the determining whether the audio signal is associated with the user comprises:
determining respective distances between a source of the audio signal and each of the plurality of microphones, and
determining, based on the respective distances, a location of the source of the audio signal;
in accordance with a determination that the audio signal is associated with the user of the wearable head device:
determining whether the audio signal comprises a non-linguistic input; and
in accordance with a determination that the audio signal comprises the non-linguistic input, performing an action associated with the non-linguistic input; and
in accordance with a determination that the audio signal is not associated with the user of the wearable head device, forgoing determining whether the audio signal comprises the non-linguistic input.
21. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute a method comprising:
receiving, using a plurality of microphones of a wearable head device, an audio signal;
determining whether the audio signal is associated with a user of the wearable head device, wherein the determining whether the audio signal is associated with the user comprises:
determining respective distances between a source of the audio signal and each of the plurality of microphones, and
determining, based on the respective distances, a location of the source of the audio signal;
in accordance with a determination that the audio signal is associated with the user of the wearable head device:
determining whether the audio signal comprises a non-linguistic input; and
in accordance with a determination that the audio signal comprises the non-linguistic input, performing an action associated with the non-linguistic input; and
in accordance with a determination that the audio signal is not associated with the user of the wearable head device, forgoing determining whether the audio signal comprises the non-linguistic input.Join the waitlist — get patent alerts
Track US12417766B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.