US2024265912A1PendingUtilityA1

Weighted finite state transducer frameworks for conversational ai systems and applications

Assignee: NVIDIA CORPPriority: Feb 2, 2023Filed: Jul 20, 2023Published: Aug 8, 2024
Est. expiryFeb 2, 2043(~16.5 yrs left)· nominal 20-yr term from priority
G10L 15/22G10L 15/063G10L 15/16
54
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 generating a unit schema for a lattice corresponding to an auditory utterance;   generating a time schema for the lattice corresponding to the auditory utterance;   combining the unit schema and the time schema to form a combined lattice representation; and   populating the combined lattice representation with log probabilities of the auditory utterance.   
     
     
         2 . The method of  claim 1 , further comprising:
 receiving an input tensor corresponding to the auditory utterance; and   reshaping the input tensor to combine an audio dimension and a text dimension into a single value.   
     
     
         3 . The method of  claim 2 , wherein the combined lattice representation is formed from at least an epsilon adapter and a soft-alignment weighted finite state transducer. 
     
     
         4 . The method of  claim 3 , wherein the combined lattice representation is formed from at least an emissions weighted finite state transducer. 
     
     
         5 . The method of  claim 1 , wherein one or more arcs of the unit schema correspond to an epsilon emission. 
     
     
         6 . The method of  claim 5 , wherein the epsilon emission has a weight of zero. 
     
     
         7 . The method of  claim 1 , wherein labels corresponding to arcs of both the unit schema and the time schema include an input, an output, a time, a unit index, and a weight. 
     
     
         8 . A system comprising:
 at least one processor to:
 compute, using a neural network and based at least on a first frame of an auditory input, a first output value indicative of an epsilon token; 
 based at least on the epsilon token, select, as a next frame for the neural network to process after the first frame, a second frame of the auditory input; 
 compute, using the neural network and based at least on the second frame, a second output value indicative of a unit token; 
 determine the second output value indicative of the unit token is an end of an utterance of the auditory input; and 
 compute, for the utterance, an utterance value including at least the first output value and the second output value, wherein the first output value has a value of zero. 
   
     
     
         9 . The system of  claim 8 , wherein the system comprises at least one of:
 a system for performing simulation operations;   a system for performing simulation operations to test or validate autonomous machine applications;   a system for performing digital twin operations;   a system for performing light transport simulation;   a system for rendering graphical output;   a system for performing deep learning operations;   a system implemented using an edge device;   a system for generating or presenting virtual reality (VR) content;   a system for generating or presenting augmented reality (AR) content;   a system for generating or presenting mixed reality (MR) content;   a system incorporating one or more Virtual Machines (VMs);   a system for performing operations for a conversational AI application;   a system for performing operations for a generative AI application;   a system for performing operations using a language model;   a system implemented at least partially in a data center;   a system for performing hardware testing using simulation;   a system for synthetic data generation;   a collaborative content creation platform for 3D assets; or   a system implemented at least partially using cloud computing resources.   
     
     
         10 . The system of  claim 8 , wherein the neural network includes a recurrent neural network transducer (RNN-T). 
     
     
         11 . The system of  claim 8 , wherein the computing the utterance value further causes the at least one processor to:
 determine forward weights for a path from a start to the end;   determine backward weights for the path; and   determine a path cost based at least on the one or more loss functions.   
     
     
         12 . The system of  claim 8 , wherein a cost associated with the epsilon token is equal to zero. 
     
     
         13 . The system of  claim 8 , wherein the at least one processor is further to:
 generate an epsilon-adapter from a text length to emulate blank transitions;   combine the epsilon-adapter with a soft-alignment weighted finite state transducer (WFST) to form a modified WFST; and   combine the modified WFST with an emissions WFST.   
     
     
         14 . The system of  claim 13 , wherein the soft-alignment WFST is a composition of a topology WFST and a linear finite state acceptor. 
     
     
         15 . The system of  claim 13 , wherein the at least one processor is further to:
 reshape an input tensor to combine an audio dimension and a text dimension into a single value.   
     
     
         16 . A processor comprising:
 one or more processing units to perform, using a neural network (NN), one or more automatic speech recognition (ASR) operations with respect to a sequence of audio frames, wherein an output of the NN corresponding to at least one audio frame of the sequence of audio frames corresponds to an emulated blank causing an audio frame in the sequence of audio frames to be omitted from processing using the NN without affecting a loss score.   
     
     
         17 . The processor of  claim 16 , wherein the processor is comprised at least one of:
 a system for performing simulation operations;   a system for performing simulation operations to test or validate autonomous machine applications;   a system for performing digital twin operations;   a system for performing light transport simulation;   a system for rendering graphical output;   a system for performing deep learning operations;   a system implemented using an edge device;   a system for generating or presenting virtual reality (VR) content;   a system for generating or presenting augmented reality (AR) content;   a system for generating or presenting mixed reality (MR) content;   a system incorporating one or more Virtual Machines (VMs);   a system for performing operations for a conversational AI application;   a system for performing operations for a generative AI application;   a system for performing operations using a language model;   a system implemented at least partially in a data center;   a system for performing hardware testing using simulation;   a system for synthetic data generation;   a collaborative content creation platform for 3D assets; or   a system implemented at least partially using cloud computing resources.   
     
     
         18 . The processor of  claim 16 , wherein the neural network includes a recurrent neural network transducer (RNN-T). 
     
     
         19 . The processor of  claim 16 , wherein the neural network is trained using at least one path through a probability lattice. 
     
     
         20 . The processor of  claim 19 , wherein the probability lattice is formed based, at least, on a reshaped input tensor that is reduced from four dimensions to three dimensions.

Join the waitlist — get patent alerts

Track US2024265912A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.