US2017069309A1PendingUtilityA1

Enhanced speech endpointing

Assignee: GOOGLE INCPriority: Sep 3, 2015Filed: Jun 24, 2016Published: Mar 9, 2017
Est. expirySep 3, 2035(~9.1 yrs left)· nominal 20-yr term from priority
G06F 3/167G10L 15/22G10L 2015/088G10L 25/78G10L 15/04G10L 2025/783G10L 15/05G10L 15/26G10L 15/1815
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for receiving audio data including an utterance, obtaining context data that indicates one or more expected speech recognition results, determining an expected speech recognition result based on the context data, receiving an intermediate speech recognition result generated by a speech recognition engine, comparing the intermediate speech recognition result to the expected speech recognition result for the audio data based on the context data, determining whether the intermediate speech recognition result corresponds to the expected speech recognition result for the audio data based on the context data, and setting an end of speech condition and providing a final speech recognition result in response to determining the intermediate speech recognition result matches the expected speech recognition result, the final speech recognition result including the one or more expected speech recognition results indicated by the context data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . (canceled) 
     
     
         2 . A computer-implemented method comprising:
 receiving an utterance in which a user speaks one or more words that make up an initial portion of a voice command, pauses for longer than a pause duration that is associated by default with endpointing utterances, then completes the voice command by speaking one or more additional words; and   submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words.   
     
     
         3 . The method of  claim 2 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes the transcription of the one or more words that make up the initial portion without the one or more additional words. 
     
     
         4 . The method of  claim 2 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes only the transcription of the one or more words that make up the initial portion. 
     
     
         5 . The method of  claim 2 , comprising:
 after the user pauses for longer than the pause duration that is associated by default with endpointing utterances and before the user completes the voice command by speaking the one or more additional words, updating a user interface to include a transcription of only the one or more words that make up the initial portion of the voice command.   
     
     
         6 . The method of  claim 2 , comprising:
 updating a user interface to include the transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words,   wherein the user interface is updated after the voice command is completed in response to an indication that the one or more words that make up the initial portion of the voice command do not likely correspond to an expected speech recognition result that is based on context data, the context data indicating one or more expected speech recognition results.   
     
     
         7 . The method of  claim 2 , wherein the one or more words that make up the initial portion of the voice command correspond to a likely incomplete voice command. 
     
     
         8 . The method of  claim 6 , wherein the context data corresponds to data stored in or displayed on a client device associated with the user. 
     
     
         9 . A system comprising:
 one or more processors; and   one or more storage devices storing instructions that are operable, when executed by the one or more processors, to cause the one or more processors to perform operations comprising:
 receiving an utterance in which a user speaks one or more words that make up an initial portion of a voice command, pauses for longer than a pause duration that is associated by default with endpointing utterances, then completes the voice command by speaking one or more additional words; and 
 submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words. 
   
     
     
         10 . The system of  claim 9 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes the transcription of the one or more words that make up the initial portion without the one or more additional words. 
     
     
         11 . The system of  claim 9 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes only the transcription of the one or more words that make up the initial portion. 
     
     
         12 . The system of  claim 9 , the operations comprising:
 after the user pauses for longer than the pause duration that is associated by default with endpointing utterances and before the user completes the voice command by speaking the one or more additional words, updating a user interface to include a transcription of only the one or more words that make up the initial portion of the voice command.   
     
     
         13 . The system of  claim 9 , the operations comprising:
 updating a user interface to include the transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words,   wherein the user interface is updated after the voice command is completed in response to an indication that the one or more words that make up the initial portion of the voice command do not likely correspond to an expected speech recognition result that is based on context data, the context data indicating one or more expected speech recognition results.   
     
     
         14 . The system of  claim 9 , wherein the one or more words that make up the initial portion of the voice command correspond to a likely incomplete voice command. 
     
     
         15 . The system of  claim 9 , the operations comprising turning off an audio input device into which the utterance was made. 
     
     
         16 . The method of  claim 14 , wherein the context data corresponds to data stored in or displayed on a client device associated with the user. 
     
     
         17 . A non-transitory, computer-readable medium storing instructions executable by one or more processors which, upon such execution, cause the one or more processors to perform operations comprising:
 receiving an utterance in which a user speaks one or more words that make up an initial portion of a voice command, pauses for longer than a pause duration that is associated by default with endpointing utterances, then completes the voice command by speaking one or more additional words; and   submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words.   
     
     
         18 . The computer-readable medium of  claim 17 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes the transcription of the one or more words that make up the initial portion without the one or more additional words. 
     
     
         19 . The computer-readable medium of  claim 17 , wherein submitting the voice command that includes a transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words occurs without submitting a voice command that includes only the transcription of the one or more words that make up the initial portion. 
     
     
         20 . The computer-readable medium of  claim 17 , the operations comprising:
 after the user pauses for longer than the pause duration that is associated by default with endpointing utterances and before the user completes the voice command by speaking the one or more additional words, updating a user interface to include a transcription of only the one or more words that make up the initial portion of the voice command.   
     
     
         21 . The computer-readable medium of  claim 17 , the operations comprising:
 updating a user interface to include the transcription of the one or more words that make up the initial portion of the voice command and the one or more additional words,   wherein the user interface is updated after the voice command is completed in response to an indication that the one or more words that make up the initial portion of the voice command do not likely correspond to an expected speech recognition result that is based on context data, the context data indicating one or more expected speech recognition results.

Join the waitlist — get patent alerts

Track US2017069309A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.