US2017270914A1PendingUtilityA1
Method and apparatus for evaluating trigger phrase enrollment
Assignee: Google Technology Holdings LLCPriority: Jul 31, 2013Filed: Jun 5, 2017Published: Sep 21, 2017
Est. expiryJul 31, 2033(~7 yrs left)· nominal 20-yr term from priority
G10L 15/063G10L 15/1807G10L 21/0264G10L 25/84G10L 2015/088G10L 15/20
55
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
An electronic device includes a microphone that receives an audio signal that includes a spoken trigger phrase, and a processor that is electrically coupled to the microphone. The processor measures characteristics of the audio signal, and determines, based on the measured characteristics, whether the spoken trigger phrase is acceptable for trigger phrase model training. If the spoken trigger phrase is determined not to be acceptable for trigger phrase model training, the processor rejects the trigger phrase for trigger phrase model training.
Claims
exact text as granted — not AI-modified1 - 9 . (canceled)
10 . A computer-implemented method comprising:
during a hotword enrollment process, prompting a user to speak a candidate hotword, and receiving audio data corresponding to the user speaking the candidate hotword; and in response to determining that a length of the spoken candidate hotword satisfies a threshold, prompting the user to speak the candidate hotword again.
11 . The computer-implemented method of claim 10 , wherein prompting the user to speak the candidate hotword again is in response to rejecting the candidate hotword spoken in a first attempt during the hotword enrollment process.
12 . The computer-implemented method of claim 10 , comprising:
identifying, by a computing device, audio characteristics of the spoken candidate hotword and audio characteristics of the background noise for each frame in the received audio data; comparing, by the computing device, the identified audio characteristics of the spoken candidate hotword to predetermined threshold values associated with unacceptable values for trigger phrase model training; and determining, by the computing device, a voice activity detection flag for each of the frames in the received audio data in response to comparing the identified audio characteristics of the spoken candidate hotword to predetermined threshold values.
13 . The computer-implemented method of claim 12 , wherein determining the voice activity detection flag for each of the frames in the received audio data comprises:
generating, by the computing device, an accept enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being less than the predetermined threshold values; and generating, by the computing device, a reject enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being greater than the predetermined threshold values.
14 . The computer-implemented method of claim 13 , comprising:
determining, by the computing device, the length of the spoken candidate hotword in the received audio data comprises determining a number of frames in the received audio signal that obtains the accept enrollment flag.
15 . The computer-implemented method of claim 14 , comprising:
comparing, by the computing device, the length of the spoken candidate hotword to a lower phrase length threshold and to a higher phrase length threshold in response to determining the number of frames in the received audio signal that obtain the accept enrollment flag; and in response to comparing the length of the spoken candidate hotword to the lower phrase length threshold and to the higher phrase threshold, prompting, by the computing device, the user to speak the candidate hotword in a second attempt, in which the length of the spoken candidate hotword is less than the lower phrase length threshold or greater than the higher phrase threshold.
16 . The computer-implemented method of claim 15 , wherein the lower phrase length threshold is less than 70 frames in which the accept enrollment flag is found, the higher phrase length threshold is greater than 180 frames in which the accept enrollment flag is found, and each frame is 10 milliseconds in duration.
17 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: during a hotword enrollment process, prompting a user to speak a candidate hotword, and receiving audio data corresponding to the user speaking the candidate hotword; and in response to determining that a length of the spoken candidate hotword satisfies a threshold, prompting the user to speak the candidate hotword again.
18 . The system of claim 17 , wherein prompting the user to speak the candidate hotword again is in response to rejecting the candidate hotword spoken in a first attempt during the hotword enrollment process.
19 . The system of claim 17 , wherein the operations further comprise:
identifying, by a computing device, audio characteristics of the spoken candidate hotword and audio characteristics of the background noise for each frame in the received audio data; comparing, by the computing device, the identified audio characteristics of the spoken candidate hotword to predetermined threshold values associated with unacceptable values for trigger phrase model training; and determining, by the computing device, a voice activity detection flag for each of the frames in the received audio data in response to comparing the identified audio characteristics of the spoken candidate hotword to predetermined threshold values.
20 . The system of claim 19 , wherein determining the voice activation detection flag for each of the frames in the received audio data the operations further comprise:
generating, by the computing device, an accept enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being less than the predetermined threshold values; and generating, by the computing device, a reject enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being greater than the predetermined threshold values.
21 . The system of claim 20 , wherein the operations further comprise:
determining, by the computing device, the length of the spoken candidate hotword in the received audio data comprises determining a number of frames in the received audio signal that obtains the accept enrollment flag.
22 . The system of claim 21 , wherein the operations further comprise:
comparing, by the computing device, the length of the spoken candidate hotword to a lower phrase length threshold and to a higher phrase length threshold in response to determining the number of frames in the received audio signal that obtain the accept enrollment flag; and in response to comparing the length of the spoken candidate hotword to the lower phrase length threshold and to the higher phrase threshold, prompting, by the computing device, the user to speak the candidate hotword in a second attempt, in which the length of the spoken candidate hotword is less than the lower phrase length threshold or greater than the higher phrase threshold.
23 . The system of claim 22 , wherein the lower phrase length threshold is less than 70 frames in which the accept enrollment flag is found, the higher phrase length threshold is greater than 180 frames in which the accept enrollment flag is found, and each frame is 10 milliseconds in duration.
24 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: during a hotword enrollment process, prompting a user to speak a candidate hotword, and receiving audio data corresponding to the user speaking the candidate hotword; and in response to determining that a length of the spoken candidate hotword satisfies a threshold, prompting the user to speak the candidate hotword again.
25 . The computer-readable medium of claim 24 , wherein prompting the user to speak the candidate hotword again is in response to rejecting the candidate hotword spoken in a first attempt during the hotword enrollment process.
26 . The computer-readable medium of claim 24 , wherein the operations comprise:
identifying, by a computing device, audio characteristics of the spoken candidate hotword and audio characteristics of the background noise for each frame in the received audio data; comparing, by the computing device, the identified audio characteristics of the spoken candidate hotword to predetermined threshold values associated with unacceptable values for trigger phrase model training; and determining, by the computing device, a voice activity detection flag for each of the frames in the received audio data in response to comparing the identified audio characteristics of the spoken candidate hotword to predetermined threshold values.
27 . The computer-readable medium of claim 26 , wherein determining the voice activity detection flag for each of the frames in the received audio data the operations comprise:
generating, by the computing device, an accept enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being less than the predetermined threshold values; and generating, by the computing device, a reject enrollment flag in response to the identified audio characteristics of the spoken candidate hotword being greater than the predetermined threshold values.
28 . The computer-readable medium of claim 27 , wherein the operations comprise:
determining, by the computing device, the length of the spoken candidate hotword in the received audio data comprises determining a number of frames in the received audio signal that obtains the accept enrollment flag.
29 . The computer-readable medium of claim 28 , wherein the operations comprise:
comparing, by the computing device, the length of the spoken candidate hotword to a lower phrase length threshold and to a higher phrase length threshold in response to determining the number of frames in the received audio signal that obtain the accept enrollment flag; and in response to comparing the length of the spoken candidate hotword to the lower phrase length threshold and to the higher phrase threshold, prompting, by the computing device, the user to speak the candidate hotword in a second attempt, in which the length of the spoken candidate hotword is less than the lower phrase length threshold or greater than the higher phrase threshold.Join the waitlist — get patent alerts
Track US2017270914A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.