US2021050003A1PendingUtilityA1

Custom Wake Phrase Training

Assignee: ZAHEER SAMEER SYEDPriority: Aug 15, 2019Filed: Aug 15, 2019Published: Feb 18, 2021
Est. expiryAug 15, 2039(~13 yrs left)· nominal 20-yr term from priority
G10L 15/063G10L 15/16G10L 15/22G10L 2015/0631G10L 2015/0638G06F 40/289G10L 2015/223G10L 2015/088G06F 8/41
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

There is provided a computing system for training a model for spotting of a custom wake phrase, the system comprising a processor and a memory in communication with the processor, the memory storing instructions that, when executed by the processor, configure the computing system to: receive a request for training the model for a phrase spotter executable for spotting the custom wake phrase within audio input. The computing system is configured to: responsive to receiving the request, receive, at a user interface of the computing system: an input positive audio sample corresponding to spoken audio of the custom wake phrase; and train using the input positive audio sample, the model for the phrase spotter executable for a wake phrase recognition subsystem that, when deployed on a voice enabled computing device, recognizes subsequent audio input instances of the custom wake phrase based upon the training.

Claims

exact text as granted — not AI-modified
1 . A computing system for training custom phrase spotter executables for virtual assistants, the system comprising a processor and a memory in communication with the processor, the memory storing instructions that, when executed by the processor, configure the computing system to:
 receive a request for training a custom phrase spotter executable and an identification of a specific virtual assistant;   responsive to receiving the request, receive:
 one or more positive audio samples corresponding to spoken audio of a custom wake phrase; 
   train, using the positive audio samples, a model for the custom wake phrase audio; and   compile the executable, including the model, such that, when deployed on the specific virtual assistant as identified by the identification, the executable recognizes the custom wake phrase.   
     
     
         2 . The computing system of  claim 1  further configured to:
 receive text corresponding to the custom wake phrase; 
 search within a corpus of audio samples, stored on a database of the computing system, for one or more stored positive audio samples corresponding to the text; and 
 include the stored positive audio samples in the training of the model. 
 
     
     
         3 . The computing system of  claim 1  further configured to:
 receive text corresponding to the custom wake phrase; 
 apply text-to-speech (TTS) to the text to generate a synthesized positive audio sample of the custom wake phrase; and 
 include the synthesized positive audio sample in the training of the model. 
 
     
     
         4 . The computing system of  claim 1 , further configured to:
 responsive to receiving the request, receive one or more negative audio samples having audible similarities to the positive audio samples but that are not the custom wake phrase; and   include the negative audio samples in the training of the model.   
     
     
         5 . The computing system of  claim 1 , further configured to:
 search within a corpus of audio samples, stored on a database of the computing device, for one or more stored negative audio samples having audible similarities to the positive audio samples but that are not the custom wake phrase; and   include the stored negative audio sample in the training of the model as a negative sample.   
     
     
         6 . The computing system of  claim 2 , further configured to:
 generate a phoneme representation for the custom wake phrase, in dependence upon the text;   search, within the database, for a phonetically similar wake phrase sharing phonetic features with the phoneme representation and retrieve, from the database a stored positive audio sample corresponding to the phonetically similar wake phrase; and,   utilize the stored positive audio sample in the training of the model.   
     
     
         7 . The computing system of  claim 1 , further configured to:
 search, within a corpus of audio samples, stored in a database on the computing system for a stored positive audio sample having an alternate pronunciation of the custom wake phrase but that is an accurate representation of the custom wake phrase; and,   include the stored positive audio sample in the training of the model.   
     
     
         8 . The computing system of  claim 1 , wherein the positive audio samples comprise one of: a spoken input provided directly via a developer interface of the computing system; and an audio file provided to the developer interface. 
     
     
         9 . The computing system of  claim 1 , wherein the model for the custom wake phrase audio comprises a neural network receiving input audio features of the positive audio samples and outputting one or more sub phrase units for the input audio features, and the model further comprises a sub phrase unit sequence detector for detecting the custom wake phrase within the one or more output sub phrase units. 
     
     
         10 . The computing system of  claim 1 , wherein the custom wake phrase audio comprises a first wake phrase audio and a second wake phrase audio, the model comprising a neural network receiving input audio features of the positive audio samples of both the first and the second wake phrase audio and outputting one or more sub phrase units for the input audio features, and the model further comprises a first and a second sub phrase unit sequence detector each for respectively detecting a presence of either one of the first and the second wake phrase audio within the one or more output sub phrase units. 
     
     
         11 . The computing system of  claim 1 , wherein the custom wake phrase audio comprises a plurality of wake phrase audio, the model comprising a recurrent neural network receiving input audio features of the positive audio samples of each of the plurality of wake phrase audio and outputting one or more hidden audio features, the model configured to detect a presence of any of the plurality of wake phrase audio. 
     
     
         12 . A computer implemented method for training a custom phrase spotter executable, the method comprising:
 receiving a request for training a custom phrase spotter executable;   receiving one or more positive audio samples corresponding to spoken audio of a custom wake phrase;   training, using the positive audio samples, a model for the custom wake phrase; and   compiling the executable, including the model, such that, when deployed for a virtual assistant, the executable recognizes the custom wake phrase.   
     
     
         13 . The method of  claim 12  further comprising:
 receiving text corresponding to the custom wake phrase; 
 searching within a corpus of audio samples for stored positive audio samples corresponding to the text; and 
 including the stored positive audio samples in the training of the model. 
 
     
     
         14 . The method of  claim 12  further comprising:
 receiving text corresponding to the custom wake phrase; 
 applying text-to-speech to the text to generate a synthesized positive audio sample of the custom wake phrase; and 
 including the synthesized positive audio sample in the training of the model. 
 
     
     
         15 . The method of  claim 12 , further comprising:
 receiving one or more negative audio samples having audible similarities to the positive audio samples but that are not the custom wake phrase; and   including the negative audio samples in the training of the model.   
     
     
         16 . The method of  claim 12 , further comprising:
 searching within a corpus of audio samples for negative audio samples having audible similarities to the positive audio samples but that are not the custom wake phrase; and   including the stored negative audio samples in the training of the model.   
     
     
         17 . The method of  claim 12 , further comprising:
 searching within a corpus of audio samples for stored positive audio samples acoustically similar to the received one or more positive audio samples; and   including the stored positive audio samples in the training of the model.   
     
     
         18 . The method of  claim 12 , wherein the model for the custom wake phrase comprises:
 a neural network receiving input audio features of the positive audio samples and outputting one or more sub phrase units for the input audio features; and   a sub phrase unit sequence detector for detecting the custom wake phrase within the one or more output sub phrase units.   
     
     
         19 . The method of  claim 12 , wherein the positive audio samples comprise audio samples of a first wake phrase and audio samples of a second wake phrase, the model comprising:
 a neural network receiving input audio features of the positive audio samples of both the first and the second wake phrase audio and outputting one or more sub phrase units for the input audio features; and   a first and a second sub phrase unit sequence detector each for respectively detecting a presence of either one of the first and the second wake phrase audio within the one or more output sub phrase units.   
     
     
         20 . The method of  claim 12 , wherein the positive audio samples comprise audio samples of a plurality of wake phrases, the model comprising a recurrent neural network configured to audio of any of the plurality of wake phrases. 
     
     
         21 . A non-transitory computer readable medium storing code for a software development kit (SDK) for training a custom phrase spotter executable for a virtual assistant, the code is executable by a processor and that, when executed by the processor, causes the SDK to:
 receive a request for training a custom phrase spotter executable;   receive one or more positive audio samples corresponding to spoken audio of a custom wake phrase;   train, using the positive audio samples, a model for the custom phrase spotter executable; and   compile the phrase spotter executable, including the model, such that, when deployed on the virtual assistant, the executable recognizes the custom wake phrase.   
     
     
         22 . A computing system for training custom phrase spotter executables for virtual assistants, the system comprising a processor and a memory in communication with the processor, the memory storing instructions that, when executed by the processor, configure the computing system to:
 receive a request for training a custom phrase spotter executable and an identification of a specific virtual assistant;   responsive to receiving the request, receive:
 text corresponding to the custom wake phrase; 
   search within a corpus of audio samples, stored on a database of the computing system, for one or more stored positive audio samples corresponding to the text; and   train, using the positive audio samples, a model for the custom wake phrase audio; and   compile the executable, including the model, such that, when deployed on the specific virtual assistant as identified by the identification, the executable recognizes the custom wake phrase.   
     
     
         23 . The computing system of  claim 22  further configured to:
 apply text-to-speech (TTS) to the text to generate a synthesized positive audio sample of the custom wake phrase; and 
 include the synthesized positive audio sample in the training of the model. 
 
     
     
         24 . The computing system of  claim 22 , further configured to:
 search within a corpus of audio samples, stored on a database of the computing system, for one or more negative audio samples having audible similarities to the positive audio samples but that are not the custom wake phrase; and   include the negative audio samples in the training of the model.   
     
     
         25 . The computing system of  claim 22 , further configured to:
 receive input, from a developer indicating a modification request to modify the model;   responsive to the modification request, search within the corpus of audio samples, stored on the database of the computing system, for one or more additional stored positive audio samples corresponding to an additional custom wake phrase; and   include the additional positive audio sample in the training of the model.   
     
     
         26 . The computing system of  claim 22  further configured to:
 subsequent to the deploying of the model, receive feedback from a developer, indicative of the model for the phrase spotter executable recognizing incorrect audio samples as the custom wake phrase; 
 dynamically re-train the model by including the incorrect audio samples as negative samples to generate an updated model. 
 
     
     
         27 . The computing system of  claim 22 , wherein the model for the custom wake phrase audio comprises a neural network receiving input audio features of the positive audio samples and outputting one or more sub phrase units for the input audio features, the model further comprises a sub phrase unit sequence detector for detecting the custom wake phrase within the one or more output sub phrase units. 
     
     
         28 . The computing system of  claim 22 , wherein the custom wake phrase audio comprises a plurality of wake phrase audio, the model comprising a recurrent neural network receiving input audio features of the positive audio samples of each of the plurality of wake phrase audio and outputting one or more hidden audio features, the model configured to detect a presence of any of the plurality of wake phrase audio.

Join the waitlist — get patent alerts

Track US2021050003A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.