US11887580B2ActiveUtilityA1

Dynamic system response configuration

Assignee: AMAZON TECH INCPriority: Dec 10, 2020Filed: Jan 4, 2023Granted: Jan 30, 2024
Est. expiryDec 10, 2040(~14.4 yrs left)· nominal 20-yr term from priority
G10L 13/047G10L 13/086G10L 15/18G10L 15/22G06F 40/30G06F 3/167G10L 13/033G10L 25/63G10L 15/26G06N 7/01G06N 3/045G06N 3/0442G06N 3/09
92
PatentIndex Score
5
Cited by
6
References
20
Claims

Abstract

A natural language processing system may select a synthesized speech quality using user profile data. The system may receive a natural language input and determine responsive output data. The system may, based at least in part on user profile data associated with the input, determine response configuration data corresponding to a quality of synthesized speech. The system may then determine further output data for presentation using the responsive output data and response configuration data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A computer-implemented method comprising:
 receiving input data corresponding to a natural language input; 
 determining user profile data associated with the natural language input; 
 determining, based at least in part on the user profile data, first system response configuration data corresponding to the natural language input, the first system response configuration data corresponding to a first quality of synthesized speech; 
 determining, using the input data, first output data responsive to the natural language input; 
 determining, based at least in part on the first output data and the first system response configuration data, second output data; and 
 causing the second output data to be presented in response to the natural language input. 
 
     
     
       2. The computer-implemented method of  claim 1 , further comprising:
 processing the first output data to determine third output data; 
 determining the third output data does not correspond to the first system response configuration data; and 
 based at least in part on the third output data not corresponding to the first system response configuration data, selecting the second output data to be presented. 
 
     
     
       3. The computer-implemented method of  claim 1 , further comprising:
 performing speech processing using the input data to determine speech processing results data; 
 processing the speech processing results data using a skill component to determine the first output data; and 
 performing speech synthesis processing using first data representing the first output data to determine the second output data, the second output data corresponding to synthesized speech having the first quality. 
 
     
     
       4. The computer-implemented method of  claim 3 , further comprising:
 selecting a speech synthesis voice profile based at least in part on the user profile data, wherein the speech synthesis voice profile corresponds to the first system response configuration data. 
 
     
     
       5. The computer-implemented method of  claim 1 , further comprising:
 processing the input data to determine sentiment data corresponding to the natural language input; and 
 processing the sentiment data and the user profile data to determine the first system response configuration data. 
 
     
     
       6. The computer-implemented method of  claim 5 , wherein the input data comprises audio data and wherein processing the input data to determine sentiment data comprises processing the audio data to determine the sentiment data. 
     
     
       7. The computer-implemented method of  claim 1 , further comprising:
 determining, based at least in part on the user profile data, second system response configuration data corresponding to the natural language input, the second system response configuration data corresponding to selection of words for a system response; 
 performing natural language generation using the second system response configuration data to determine the first output data; and 
 performing speech synthesis processing using the first output data to determine the second output data, the second output data corresponding to synthesized speech having the first quality. 
 
     
     
       8. The computer-implemented method of  claim 1 , further comprising:
 determining dialog data corresponding to a plurality of prior inputs associated with the user profile data; and 
 using the dialog data to select a first system response configuration profile from a plurality of system response configuration profiles, the first system response configuration profile corresponding to the first system response configuration data. 
 
     
     
       9. The computer-implemented method of  claim 1 , further comprising:
 using a dialog management component, the first output data, and the first system response configuration data to determine the second output data. 
 
     
     
       10. The computer-implemented method of  claim 1 , further comprising:
 determining first data representing a state of a user corresponding to the natural language input; and 
 processing the first data and the user profile data to determine the first system response configuration data. 
 
     
     
       11. A system comprising:
 at least one processor; and 
 at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
 receive input data corresponding to a natural language input; 
 determine user profile data associated with the natural language input; 
 determine, based at least in part on the user profile data, first system response configuration data corresponding to the natural language input, the first system response configuration data corresponding to selection of words for a system response; 
 determine, using the input data, first output data responsive to the natural language input; 
 determine, based at least in part on the first output data and the first system response configuration data, second output data; and 
 cause the second output data to be presented in response to the natural language input. 
 
 
     
     
       12. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the first output data to determine third output data; 
 determine the third output data does not correspond to the first system response configuration data; and 
 based at least in part on the third output data not corresponding to the first system response configuration data, select the second output data to be presented. 
 
     
     
       13. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 perform speech processing using the input data to determine speech processing results data; 
 process the speech processing results data using a skill component to determine the first output data; and 
 perform natural language generation using the first system response configuration data and first data representing the first output data to determine the second output data. 
 
     
     
       14. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the input data to determine sentiment data corresponding to the natural language input; and 
 process the sentiment data and the user profile data to determine the first system response configuration data. 
 
     
     
       15. The system of  claim 14 , wherein the input data comprises audio data and wherein processing the input data to determine sentiment data comprises processing the audio data to determine the sentiment data. 
     
     
       16. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 determine, based at least in part on the user profile data, second system response configuration data corresponding to the natural language input, the second system response configuration data corresponding to a first quality of synthesized speech; 
 perform natural language generation using the first system response configuration data to determine first data; and 
 perform speech synthesis processing using the first data and the second system response configuration data to determine the second output data, the second output data corresponding to synthesized speech having the first quality. 
 
     
     
       17. The system of  claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 select a speech synthesis voice profile based at least in part on the user profile data, wherein the speech synthesis voice profile corresponds to the second system response configuration data. 
 
     
     
       18. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 determine dialog data corresponding to a plurality of prior inputs associated with the user profile data; and 
 use the dialog data to select a first system response configuration profile from a plurality of system response configuration profiles, the first system response configuration profile corresponding to the first system response configuration data. 
 
     
     
       19. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 use a dialog management component, the first output data, and the first system response configuration data to determine the second output data. 
 
     
     
       20. The system of  claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 determine first data representing a state of a user corresponding to the natural language input; and 
 process the first data and the user profile data to determine the first system response configuration data.

Join the waitlist — get patent alerts

Track US11887580B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.