Individualized hotword detection models
Abstract
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for presenting notifications in an enterprise system. In one aspect, a method include actions of obtaining enrollment acoustic data representing an enrollment utterance spoken by a user, obtaining a set of candidate acoustic data representing utterances spoken by other users, determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents a similarity between the enrollment acoustic data and the candidate acoustic data, selecting a subset of candidate acoustic data from the set of candidate acoustic data based at least on the similarity scores, generating a detection model based on the subset of candidate acoustic data, and providing the detection model for use in detecting an utterance spoken by the user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . (canceled)
2 . A computer-implemented method comprising:
receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation; in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation; dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and using the dynamically generated hotword detection model to detect a likely utterance of the predefined hotword in subsequently received audio data.
3 . The computer implemented method of claim 2 , comprising:
during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.
4 . The computer implemented method of claim 3 , comprising:
obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.
5 . The computer implemented method of claim 4 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users comprises:
determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.
6 . The computer implemented method of claim 5 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data comprises:
determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and determining the similarity score based on an averaging of the plurality of sub-similarity scores.
7 . The computer implemented method of claim 2 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword comprises:
training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.
8 . The computer implemented method of claim 2 , wherein the hotword detection model is based at least on the single utterance and not based on another utterance of the predefined hotword.
9 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation;
in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation;
dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and
using the dynamically generated hotword detection model to detect a likely utterance of the hotword in subsequently received audio data.
10 . The system of claim 9 , the operations further comprise:
during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.
11 . The system of claim 10 , the operations further comprise:
obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.
12 . The system of claim 11 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users the operations further comprise:
determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.
13 . The system of claim 12 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data the operations further comprise:
determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and determining the similarity score based on an averaging of the plurality of sub-similarity scores.
14 . The system of claim 9 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword the operations further comprise:
training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.
15 . The system of claim 9 , wherein the hotword detection model is based at least on the single utterance and not based on another utterance of the predefined hotword.
16 . A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
receiving audio data corresponding to a single utterance by a user of a predefined hotword, wherein the predefined hotword is pronounced by the user using a personalized, non-standard pronunciation;
in response to receiving the audio data corresponding to the single utterance by the user of the predefined hotword, downloading audio features corresponding to other users' utterances of the same, predefined hotword in a manner that is indicated as similar to the personalized, non-standard pronunciation;
dynamically generating a hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword; and
using the dynamically generated hotword detection model to detect a likely utterance of the hotword in subsequently received audio data.
17 . The computer-readable medium of claim 16 , the operations comprising:
during an enrollment process, prompting, by a client device, the user to speak the predefined hotword; and generating enrollment acoustic data using the received audio data from the user, wherein the audio data comprises the predefined hotword pronounced by the user using the personalized, non-standard pronunciation and additional one or more terms spoken by the user that trigger semantic interpretation of the one or more terms that follow the predefined hotword.
18 . The computer-readable medium of claim 17 , the operations comprising:
obtaining a set of candidate acoustic data representing utterances that were previously-spoken by the other users, wherein the other users are of a similar type of user to the user.
19 . The computer-readable medium of claim 18 , wherein obtaining the set of candidate acoustic data representing the utterances that were spoken by the other users the operations comprising:
determining, for each candidate acoustic data of the set of candidate acoustic data, a similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data.
20 . The computer-readable medium of claim 19 , wherein determining the similarity score that represents an acoustic similarity between the enrollment acoustic data and the candidate acoustic data the operations comprising:
determining a plurality of sub-similarity scores between the enrollment acoustic data and the candidate acoustic data; and determining the similarity score based on an averaging of the plurality of sub-similarity scores.
21 . The computer-readable medium of claim 16 , wherein dynamically generating the hotword detection model for the personalized, non-standard pronunciation of the predefined hotword using (i) the audio data corresponding to the single utterance by the user of the predefined hotword, and (ii) the downloaded audio features corresponding to other users' utterances of the same, predefined hotword the operations comprising:
training the hotword detection model to detect the likely utterance of the predefined hotword by the user in the subsequently received audio data corresponding to the single utterance of the predefined hotword by the user and without requiring the user to speak additional utterances of the predefined hotword.Join the waitlist — get patent alerts
Track US2017194006A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.