Methods and systems for determining a wake word
Abstract
A user device (e.g., voice assistant device, voice enabled device, smart device, computing device, etc.) may receive/detect audio content (e.g., speech, etc.) that includes a wake word and/or words similar to a wake word. The user device may require a wake word, a portion of the wake word, or words similar to the wake word to be detected prior to interacting with a user. The user device may, based on characteristics of the audio content, determine if the audio content originates from an authorized user. The user device may decrease and/or increase scrutiny applied to wake word detection based on whether audio content originates from an authorized user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving, by a computing device, audio content; based on one or more voice characteristics associated with the audio content that indicate that the audio content is associated with an authorized user, determining to lower a wake word threshold for processing the audio content; and based on a determination, using the lowered wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, causing execution of one or more operational commands associated with the audio content.
2 . The method of claim 1 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
3 . The method of claim 1 , wherein the determination that at least the portion of the audio content corresponds to the wake word or phrase comprises:
determining, based on the audio content, one or more words in the at least the portion of the audio content satisfy the lowered wake word threshold.
4 . The method of claim 1 , wherein the lowered wake word threshold is associated with a lower confidence level requirement that the audio content comprises the wake word or phrase.
5 . The method of claim 1 , wherein the lowered wake word threshold is associated with one or more authorized users comprising the authorized user and a higher wake word threshold is associated with an origin of the audio content that is not associated with the one or more authorized users.
6 . The method of claim 1 , wherein the one or more operational commands are associated with a target device and wherein causing execution of the one or more operational commands comprises sending, to the target device, the one or more operational commands.
7 . The method of claim 1 , further comprising:
receiving second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is not associated with one or more authorized users comprising the authorized user, increasing the wake word threshold, from the lowered wake word threshold, for processing the second audio content.
8 . A method comprising:
receiving, by a computing device, audio content; based on one or more voice characteristics associated with the audio content that indicate that the audio content is not associated with one or more authorized users, determining to increase a wake word threshold for processing the audio content; and based on a determination, using the increased wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, causing execution of one or more operational commands associated with the audio content.
9 . The method of claim 8 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
10 . The method of claim 8 , wherein the increased wake word threshold is associated with a higher confidence level requirement that the audio content comprises the wake word or phrase.
11 . The method of claim 8 , wherein the determination that at least the portion of the audio content corresponds to the wake word or phrase comprises:
determining, based on the audio content, one or more words in the at least the portion of the audio content satisfy the increased wake word threshold.
12 . The method of claim 8 , wherein the increased wake word threshold is greater than a lower wake word threshold associated with the one or more authorized users.
13 . The method of claim 8 , further comprising:
receiving second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is associated with one or more of the one or more authorized users, lowering the wake word threshold, from the increased wake word threshold, for processing the second audio content.
14 . The method of claim 8 , wherein the one or more operational commands are associated with a target device and wherein causing execution of the one or more operational commands comprises sending, to the target device, the one or more operational commands.
15 . One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to:
receive audio content; based on one or more voice characteristics associated with the audio content that indicate that the audio content is associated with an authorized user, determine to lower a wake word threshold for processing the audio content; and based on a determination, using the lowered wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, cause execution of one or more operational commands associated with the audio content.
16 . The one or more non-transitory computer-readable media of claim 15 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
17 . The one or more non-transitory computer-readable media of claim 15 , wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to determine that at least the portion of the audio content corresponds to the wake word or phrase, cause the at least one processor to determine, based on the audio content, one or more words in the at least the portion of the audio content satisfy the lowered wake word threshold.
18 . The one or more non-transitory computer-readable media of claim 15 , wherein the lowered wake word threshold is associated with a lower confidence level requirement that the audio content comprises the wake word or phrase.
19 . The one or more non-transitory computer-readable media of claim 15 , wherein the lowered wake word threshold is associated with one or more authorized users comprising the authorized user and a higher wake word threshold is associated with an origin of the audio content that is not associated with the one or more authorized users.
20 . The one or more non-transitory computer-readable media of claim 15 , wherein the one or more operational commands are associated with a target device and wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to cause execution of the one or more operational commands, cause the at least one processor to send, to the target device, the one or more operational commands.
21 . The one or more non-transitory computer-readable media of claim 15 , wherein the processor-executable instructions, when executed by the at least one processor, further cause the at least one processor to:
receive second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is not associated with one or more authorized users comprising the authorized user, increase the wake word threshold, from the lowered wake word threshold, for processing the second audio content.
22 . One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to:
receive audio content; based on one or more voice characteristics associated with the audio content that indicate that the audio content is not associated with one or more authorized users, determine to increase a wake word threshold for processing the audio content; and based on a determination, using the increased wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, cause execution of one or more operational commands associated with the audio content.
23 . The one or more non-transitory computer-readable media of claim 22 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
24 . The one or more non-transitory computer-readable media of claim 22 , wherein the increased wake word threshold is associated with a higher confidence level requirement that the audio content comprises the wake word or phrase.
25 . The one or more non-transitory computer-readable media of claim 22 , wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to determine that at least the portion of the audio content corresponds to the wake word or phrase, cause the at least one processor to determine, based on the audio content, one or more words in the at least the portion of the audio content satisfy the increased wake word threshold.
26 . The one or more non-transitory computer-readable media of claim 22 , wherein the increased wake word threshold is greater than a lower wake word threshold associated with the one or more authorized users.
27 . The one or more non-transitory computer-readable media of claim 22 , wherein the processor-executable instructions, when executed by the at least one processor, further cause the at least one processor to:
receive second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is associated with one or more of the one or more authorized users, lower the wake word threshold, from the increased wake word threshold, for processing the second audio content.
28 . The one or more non-transitory computer-readable media of claim 22 , wherein the one or more operational commands are associated with a target device and wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to cause execution of the one or more operational commands, cause the at least one processor to send, to the target device, the one or more operational commands.
29 . An apparatus comprising:
one or more processors; and memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to:
receive audio content;
based on one or more voice characteristics associated with the audio content that indicate that the audio content is associated with an authorized user, determine to lower a wake word threshold for processing the audio content; and
based on a determination, using the lowered wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, cause execution of one or more operational commands associated with the audio content.
30 . The apparatus of claim 29 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
31 . The apparatus of claim 29 , wherein the processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine that at least the portion of the audio content corresponds to the wake word or phrase, cause the apparatus to determine, based on the audio content, one or more words in the at least the portion of the audio content satisfy the lowered wake word threshold.
32 . The apparatus of claim 29 , wherein the lowered wake word threshold is associated with a lower confidence level requirement that the audio content comprises the wake word or phrase.
33 . The apparatus of claim 29 , wherein the lowered wake word threshold is associated with one or more authorized users comprising the authorized user and a higher wake word threshold is associated with an origin of the audio content that is not associated with the one or more authorized users.
34 . The apparatus of claim 29 , wherein the one or more operational commands are associated with a target device and wherein the processor-executable instructions that, when executed by the one or more processors, cause the apparatus to cause execution of the one or more operational commands, cause the apparatus to send, to the target device, the one or more operational commands.
35 . The apparatus of claim 29 , wherein the processor-executable instructions, when executed by the one or more processors, further cause the apparatus to:
receive second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is not associated with one or more authorized users comprising the authorized user, increase the wake word threshold, from the lowered wake word threshold, for processing the second audio content.
36 . An apparatus comprising:
one or more processors; and memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to:
receive audio content;
based on one or more voice characteristics associated with the audio content that indicate that the audio content is not associated with one or more authorized users, determine to increase a wake word threshold for processing the audio content; and
based on a determination, using the increased wake word threshold, that at least a portion of the audio content corresponds to a wake word or phrase, cause execution of one or more operational commands associated with the audio content.
37 . The apparatus of claim 36 , wherein the one or more voice characteristics comprises one or more of: a frequency, a decibel level, or a tone.
38 . The apparatus of claim 36 , wherein the increased wake word threshold is associated with a higher confidence level requirement that the audio content comprises the wake word or phrase.
39 . The apparatus of claim 36 , wherein the processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine that at least the portion of the audio content corresponds to the wake word or phrase, cause the apparatus to determine, based on the audio content, one or more words in the at least the portion of the audio content satisfy the increased wake word threshold.
40 . The apparatus of claim 36 , wherein the increased wake word threshold is greater than a lower wake word threshold associated with the one or more authorized users.
41 . The apparatus of claim 36 , wherein the processor-executable instructions, when executed by the one or more processors, further cause the apparatus to:
receive second audio content; and based on one or more second voice characteristics associated with the second audio content that indicate that the second audio content is associated with one or more of the one or more authorized users, lower the wake word threshold, from the increased wake word threshold, for processing the second audio content.
42 . The apparatus of claim 36 , wherein the one or more operational commands are associated with a target device and wherein the processor-executable instructions that, when executed by the one or more processors, cause the apparatus to cause execution of the one or more operational commands, cause the apparatus to send, to the target device, the one or more operational commands.Join the waitlist — get patent alerts
Track US2024038239A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.