US2022148590A1PendingUtilityA1

Architecture for multi-domain natural language processing

Assignee: AMAZON TECH INCPriority: Dec 19, 2012Filed: Nov 12, 2021Published: May 12, 2022
Est. expiryDec 19, 2032(~6.4 yrs left)· nominal 20-yr term from priority
G06F 40/35G10L 15/26G10L 13/08G10L 15/22G06F 40/295G06F 40/56G06F 40/284G06F 40/40
74
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Features are disclosed for processing a user utterance with respect to multiple subject matters or domains, and for selecting a likely result from a particular domain with which to respond to the utterance or otherwise take action. A user utterance may be transcribed by an automatic speech recognition (“ASR”) module, and the results may be provided to a multi-domain natural language understanding (“NLU”) engine. The multi-domain NLU engine may process the transcription(s) in multiple individual domains rather than in a single domain. In some cases, the transcription(s) may be processed in multiple individual domains in parallel or substantially simultaneously. In addition, hints may be generated based on previous user interactions and other data. The ASR module, multi-domain NLU engine, and other components of a spoken language processing system may use the hints to more efficiently process input or more accurately generate output.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A system comprising:
 computer-readable memory storing executable instructions; and   one or more processors in communication with the computer-readable memory, wherein the one or more processors are programmed by the executable instructions to:
 receive movement data representing a movement of a user device; 
 generate, using the movement data, hint data regarding a future user utterance; 
 receive, subsequent to generating the hint data, natural language data representing a user utterance; 
 determine, based at least partly on the hint data, to perform natural language understanding (“NLU”) processing on the natural language data; 
 perform the NLU processing on the natural language data to generate intent data; and 
 generate response data based at least partly on the intent data. 
   
     
     
         3 . The system of  claim 2 , wherein the one or more processors are further programmed by the executable instructions to determine, based at least partly on the hint data, that the natural language data is associated with an intent of a plurality of intents, and wherein the intent data represents the intent. 
     
     
         4 . The system of  claim 2 , wherein the one or more processors are further programmed by the executable instructions to determine, based at least partly on the hint data, that the natural language data is associated with an NLU domain of a plurality of NLU domains, and wherein the intent data represents an intent associated with the NLU domain. 
     
     
         5 . The system of  claim 4 , wherein the NLU domain is associated with intents related to at least one of: phone dialing, shopping, getting directions, playing music, or performing a search. 
     
     
         6 . The system of  claim 2 , wherein the movement of the user device comprises a vertical rise of the user device. 
     
     
         7 . The system of  claim 2 , wherein the one or more processors are further programmed by the executable instructions to:
 generate second intent data based at least partly on the natural language data, wherein the intent data represents a first intent, and wherein the second intent data represents a second intent different from the first intent; and   rank the intent data and the second intent data based at least partly on the hint data.   
     
     
         8 . The system of  claim 2 , wherein to perform the NLU processing, the one or more processors are further programmed by the executable instructions to generate NLU result data using the natural language data and an NLU subsystem of a plurality of NLU subsystems,
 wherein the NLU subsystem is associated with an NLU domain of a plurality of NLU domains, and   wherein the NLU result data represents a named entity of a plurality of named entities associated with an intent represented by the intent data.   
     
     
         9 . The system of  claim 2 , wherein the one or more processors are further programmed by the executable instructions to:
 receive audio data representing the user utterance; and   generate the natural language data using the audio data and an automatic speech recognition (“ASR”) subsystem.   
     
     
         10 . The system of  claim 2 , wherein the executable instructions to generate the response data comprise executable instructions to:
 generate response text data using a natural language generation subsystem; and   generate response audio data using the response text data and a text-to-speech subsystem.   
     
     
         11 . The system of  claim 2 , wherein one or more processors are further programmed by the executable instructions to send the response data to the user device, and wherein the user device is configured to present a response using the response data. 
     
     
         12 . A computer-implemented method comprising:
 under control of one or more computing devices configured with specific computer-executable instructions,
 receiving movement data representing a movement of a user device; 
 generating, using the movement data, hint data regarding a future user utterance; 
 receiving, subsequent to generating the hint data, natural language data representing a user utterance; 
 determining, based at least partly on the hint data, to perform natural language understanding (“NLU”) processing on the natural language data; 
 performing the NLU processing on the natural language data to generate intent data; and 
 generating response data based at least partly on the intent data. 
   
     
     
         13 . The computer-implemented method of  claim 12 , further comprising determining, based at least partly on the hint data, that the natural language data is associated with an intent of a plurality of intents, wherein the intent data represents the intent. 
     
     
         14 . The computer-implemented method of  claim 12 , further comprising determining, based at least partly on the hint data, that the natural language data is associated with an NLU domain of a plurality of NLU domains, wherein the intent data represents and intent associated with the NLU domain. 
     
     
         15 . The computer-implemented method of  claim 12 , further comprising:
 generating second intent data based at least partly on the natural language data, wherein the intent data represents a first intent, and wherein the second intent data represents a second intent different from the first intent; and   ranking the intent data and the second intent data based at least partly on the hint data.   
     
     
         16 . The computer-implemented method of  claim 16 , wherein ranking the intent data and the second intent data based at least partly on the hint data comprises adjusting a rank of at least one of the intent data or the second intent data. 
     
     
         17 . The computer-implemented method of  claim 12 , further comprising determining that the movement data represents a vertical rise of the user device. 
     
     
         18 . The computer-implemented method of  claim 12 , wherein performing the NLU processing comprises generating NLU result data using the natural language data and an NLU subsystem of a plurality of NLU subsystems,
 wherein the NLU subsystem is associated with an NLU domain of a plurality of NLU domains, and   wherein the NLU result data represents a named entity of a plurality of named entities associated with an intent represented by the intent data.   
     
     
         19 . The computer-implemented method of  claim 12 , further comprising:
 receiving audio data representing the user utterance; and   generating the natural language data using the audio data and an automatic speech recognition (“ASR”) subsystem.   
     
     
         20 . The computer-implemented method of  claim 12 , wherein generating the response data comprises:
 generating response text data using a natural language generation subsystem; and   generating response audio data using the response text data and a text-to-speech subsystem.   
     
     
         21 . The computer-implemented method of  claim 12 , further comprising sending the response data to the user device, and wherein the user device is configured to present a response using the response data.

Join the waitlist — get patent alerts

Track US2022148590A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.