Restructuring deep neural network acoustic models
Abstract
A Deep Neural Network (DNN) model used in an Automatic Speech Recognition (ASR) system is restructured. A restructured DNN model may include fewer parameters compared to the original DNN model. The restructured DNN model may include a monophone state output layer in addition to the senone output layer of the original DNN model. Singular value decomposition (SVD) can be applied to one or more weight matrices of the DNN model to reduce the size of the DNN Model. The output layer of the DNN model may be restructured to include monophone states in addition to the senones (tied triphone states) which are included in the original DNN model. When the monophone states are included in the restructured DNN model, the posteriors of monophone states are used to select a small part of senones to be evaluated.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for restructuring a Deep Neural Network (DNN) model, comprising:
accessing a DNN model that includes weight matrices and layers comprising: an input layer; hidden layers; and an output layer; reducing a sparseness of a weight matrix in the DNN model; and restructuring the DNN model with the weight matrix reduced in sparseness.
2 . The method of claim 1 , wherein reducing the sparseness of the weight matrix comprises applying a Singular Value Decomposition (SVD) to the weight matrix.
3 . The method of claim 1 , wherein applying the SVD to the weight matrix comprises decomposing the weight matrix into two matrices having smaller dimensions as compared to a size of the dimension of the weight matrix before applying the SVD to the weight matrix.
4 . The method of claim 1 , wherein restructuring the DNN model with the weight matrix reduced in sparseness comprises splitting one of the layers in the DNN model into a first layer and a second layer.
5 . The method of claim 1 , wherein reducing the sparseness of the weight matrix in the DNN model comprises reducing the sparseness of each weight matrix in the DNN model.
6 . The method of claim 1 , wherein the output layer comprises a senone output layer and a monophone state output layer.
7 . The method of claim 1 , further comprising training the output layer of the DNN to use a monophone state.
8 . The method of claim 1 , further comprising tuning the restructured model using a back-propagation method.
9 . A computer-readable storage medium storing computer-executable instructions that when executed using a processor perform actions, comprising:
accessing a restructured Deep Neural Network (DNN) model that includes one or more weight matrices reduced in size as compared to the corresponding one or more weight matrices in an original DNN model and layers comprising: an input layer; hidden layers; and an output layer; and using the restructured DNN model to recognize received utterances.
10 . The computer-readable storage medium of claim 9 , wherein a sparseness of the one or more weight matrices are reduced in size as compared to a sparseness of the one or more weight matrices of the original DNN.
11 . The computer-readable storage medium of claim 9 , wherein the output layer of the restructured DNN comprises a monophone state output layer and a senone output layer.
12 . The computer-readable storage medium of claim 9 , further comprising using posteriors of monophone states to select senones to be evaluated to reduce the number of calculations in the senone output layer.
13 . The computer-readable storage medium of claim 9 , wherein the one or more weight matrices of the restructured DNN comprises two matrices having smaller dimensions as compared to a size of the dimension of the weight matrix in the original DNN model.
14 . The computer-readable storage medium of claim 9 , further comprising tuning the restructured DNN model using a back-propagation method.
15 . A system for restructuring a Deep Neural Network (DNN) model, comprising:
a processor and memory; an operating environment executing using the processor; and a model manager that is configured to perform actions comprising:
accessing a DNN model that includes weight matrices and layers comprising: an input layer; hidden layers; and an output layer;
reducing a sparseness of weight matrices in the DNN model by removing weight parameters that are below a threshold value; and
restructuring the DNN model with the weight matrix reduced in sparseness.
16 . The system of claim 15 , wherein reducing the sparseness of the weight matrix comprises applying a Singular Value Decomposition (SVD) to the weight matrix.
17 . The system of claim 15 , wherein applying the SVD to the weight matrix comprises decomposing the weight matrix into two matrices having smaller dimensions as compared to a size of the dimension of the weight matrix before applying the SVD to the weigh matrix.
18 . The system of claim 15 , wherein restructuring the DNN model with the weight matrix reduced in sparseness comprises splitting one of the layers in the DNN model into a first layer and a second layer.
19 . The system of claim 15 , wherein the output layer comprises a senone output layer and a monophone state output layer.
20 . The system of claim 15 , further comprising training the output layer of the DNN to use a monophone state.Join the waitlist — get patent alerts
Track US2017337918A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.