US2022262367A1PendingUtilityA1

Voice Query QoS based on Client-Computed Content Metadata

Assignee: GOOGLE LLCPriority: Feb 6, 2019Filed: May 2, 2022Published: Aug 18, 2022
Est. expiryFeb 6, 2039(~12.5 yrs left)· nominal 20-yr term from priority
G10L 15/22G06F 3/167G10L 2015/088G10L 15/08G10L 15/30H04L 65/80G10L 2015/226H04L 67/568H04L 67/61G06F 16/63G10L 2015/225G06F 3/16G10L 15/28G10L 17/00G10L 25/60G10L 15/26
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes receiving an automated speech recognition (ASR) request from a user device that includes a speech input captured by the user device and content metadata associated with the speech input. The content metadata is generated by the user device. The method also includes determining a priority score for the ASR request based on the content metadata associated with the speech input and caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score. The pending ASR requests in the pre-processing backlog are ranked in order of the priority scores. The method also includes providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed before pending ASR requests associated with lower priority scores.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method that when executed on data processing hardware of a user device causes the user device to perform operations comprising:
 generating an automated speech recognition (ASR) request, the ASR request comprising:
 a speech input captured by the user device that includes a voice query; and 
 content metadata associated with the speech input, the content metadata generated by the user device; 
   receiving on-device processing instructions from a server-side query processing stack;   determining whether the server-side query processing stack is overloaded; and   when the server-side query processing stack is overloaded, executing the on-device processing instructions to identify one or more criteria for locally processing at least a portion of speech input on-device.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein the content metadata associated with the speech input represents a likelihood that the corresponding ASR request will be successfully processed by the server-side query processing stack. 
     
     
         3 . The computer-implemented method of  claim 1 , wherein the content metadata associated with the speech input represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device. 
     
     
         4 . The computer-implemented method of  claim 1 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a login indicator indicating whether or not a user associated with the user device is logged in to the user device;   a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device; or   a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source.   
     
     
         5 . The computer-implemented method of  claim 1 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;   an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;   an audio signal score of the speech input; or a spatial-localization score indicating a distance and position of a user relative to the user device.   
     
     
         6 . The computer-implemented method of  claim 1 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a transcription of the speech input generated by an on-device ASR module residing on the user device;   a user device behavior signal indicating a current behavior of the user device; or an environmental condition signal indicating current environmental conditions relative to the user device.   
     
     
         7 . The computer-implemented method of  claim 1 , wherein determining whether the server-side query processing stack is overloaded is based on at least one of:
 historical data associated with previous ASR requests communicated by the user device to the server-side query processing stack;   a schedule of past and/or predicted overload conditions at the server-side query processing stack; or   receiving an overload condition status notification from the server-side query processing stack on the fly indicating a present overload condition at the server-side query processing stack.   
     
     
         8 . The computer-implemented method of  claim 1 , wherein executing the on-device processing instructions further comprises:
 transcribing, by the data processing hardware, the speech input using a local ASR module residing on the user device;   interpreting the transcription of the speech input to determine a voice query corresponding to the speech input;   determining whether the user device can execute an action associated with the voice query corresponding to the speech input; and   executing the action associated with the voice query when the user device is able to execute the action.   
     
     
         9 . The computer-implemented method of  claim 1 , wherein executing the on-device processing instructions to identify the one or more criteria comprises executing the on-device processing instructions to identify one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the server-side query processing stack. 
     
     
         10 . The computer-implemented method of  claim 9 , wherein the operations further comprise dropping ASR request when at least one of the thresholds are dissatisfied. 
     
     
         11 . A system comprising:
 data processing hardware of a user device; and   memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:   generating an automated speech recognition (ASR) request, the ASR request comprising:
 a speech input captured by the user device that includes a voice query; and 
 content metadata associated with the speech input, the content metadata generated by the user device; 
   receiving on-device processing instructions from a server-side query processing stack;   determining whether the server-side query processing stack is overloaded; and
 when the server-side query processing stack is overloaded, executing the on-device processing instructions to identify one or more criteria for locally processing at least a portion of speech input on-device. 
   
     
     
         12 . The system of  claim 11 , wherein the content metadata associated with the speech input represents a likelihood that the corresponding ASR request will be successfully processed by the server-side query processing stack. 
     
     
         13 . The system of  claim 11 , wherein the content metadata associated with the speech input represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device. 
     
     
         14 . The system of  claim 11 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a login indicator indicating whether or not a user associated with the user device is logged in to the user device;   a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device; or   a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source.   
     
     
         15 . The system of  claim 11 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;   an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;   an audio signal score of the speech input; or   a spatial-localization score indicating a distance and position of a user relative to the user device.   
     
     
         16 . The system of  claim 11 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
 a transcription of the speech input generated by an on-device ASR module residing on the user device;   a user device behavior signal indicating a current behavior of the user device; or an environmental condition signal indicating current environmental conditions relative to the user device.   
     
     
         17 . The system of  claim 11 , wherein determining whether the server-side query processing stack is overloaded is based on at least one of:
 historical data associated with previous ASR requests communicated by the user device to the server-side query processing stack;   a schedule of past and/or predicted overload conditions at the server-side query processing stack; or   receiving an overload condition status notification from the server-side query processing stack on the fly indicating a present overload condition at the server-side query processing stack.   
     
     
         18 . The system of  claim 11 , wherein executing the on-device processing instructions further comprises:
 transcribing, by the data processing hardware, the speech input using a local ASR module residing on the user device;   interpreting the transcription of the speech input to determine a voice query corresponding to the speech input;   determining whether the user device can execute an action associated with the voice query corresponding to the speech input; and   executing the action associated with the voice query when the user device is able to execute the action.   
     
     
         19 . The system of  claim 11 , wherein executing the on-device processing instructions to identify the one or more criteria comprises executing the on-device processing instructions to identify one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the server-side query processing stack. 
     
     
         20 . The system of  claim 19 , wherein the operations further comprise dropping ASR request when at least one of the thresholds are dissatisfied.

Join the waitlist — get patent alerts

Track US2022262367A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.