Training a Neural Network Using Small Training Datasets
Abstract
Training datasets are determined for training neural networks. An input dataset comprising a plurality of samples is provided as training dataset to the neural network. Vector representations of samples of the input dataset are obtained from a hidden layer of the neural network. The samples are clustered using the vector representation. The samples are scored based on a metric that indicates the similarity of the sample to its cluster. A subset of samples is determined by excluding samples that have high similarity with their clusters. The subset of samples is labelled and used for training the neural network.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method comprising:
receiving a target sample containing one or more features; and inputting the target sample to a neural network to determine at least one of the features, the neural network trained by a training process, the neural network comprising a plurality of layers of nodes, the plurality of layers of nodes including a hidden layer of nodes, the training process comprising:
receiving a first training dataset comprising a plurality of training samples;
executing the neural network to process the training samples;
for each training sample, receiving a vector representing the training sample, the vector generated by nodes of the hidden layer of nodes;
selecting a subset of training samples based on distances between vectors representing the training samples;
for each selected training sample, associating the selected training sample with a label representing an expected result that should be output by the neural network for the selected training sample; and
providing the selected training samples of the subset as a second training dataset for retraining the neural network.
2 . The method of claim 1 , wherein selecting the subset of training samples comprises:
clustering the training samples to generate a plurality of clusters, the clustering associating each training sample with a cluster, the clustering based on distances between the vectors representing the training samples; for each training sample, determining a score indicative of similarity of the training sample with the cluster associated with the training sample; and excluding one or more training samples based on the scores to determine the subset of training samples.
3 . The method of claim 2 , wherein the one or more training samples are excluded responsive to each of the one or more training samples having a score indicating high similarity of the training sample to the cluster associated with the training sample.
4 . The method of claim 1 , wherein the plurality of layers of nodes represents a sequence of layers of nodes comprising a plurality of hidden layers, wherein the hidden layer of claim 1 is the last hidden layer from the plurality of hidden layers.
5 . The method of claim 1 , wherein the neural network further comprises an output layer of nodes for outputting a result of executing the neural network, wherein the hidden layer represents nodes that provide inputs to the output layer of nodes.
6 . The method of claim 1 , wherein the training process further comprises:
receiving a second plurality of training samples; executing the neural network to process the second plurality of training samples; for each of the training sample in the second plurality of training samples, receiving a second vector representing the training sample, the second vector generated by the nodes of the hidden layer of nodes; selecting a second subset of training samples from the second plurality of training samples based on distances between the vectors representing the training samples in the second plurality of training samples; and providing the second subset as a third training dataset for retraining the neural network.
7 . The method of claim 6 , wherein the training process is repeatedly performed until an aggregate measure of difference between a result outputted by the neural network and the expected result is below a threshold value.
8 . The method of claim 1 , wherein the neural network represents a deep learning model.
9 . The method of claim 1 , wherein the target sample is a natural language description of an input image.
10 . The method of claim 1 , wherein the neural network is an autoencoder configured to receive an input and regenerate the input as output of the neural network.
11 . A non-transitory computer readable medium configured to store computer code comprising instructions, the instructions, when executed by one or more processors, cause the one or more processors to perform steps comprising:
receiving a target sample containing one or more features; and inputting the target sample to a neural network to determine at least one of the features, the neural network trained by a training process, the neural network comprising a plurality of layers of nodes, the plurality of layers of nodes including a hidden layer of nodes, the training process comprising:
receiving a first training dataset comprising a plurality of training samples;
executing the neural network to process the training samples;
for each training sample, receiving a vector representing the training sample, the vector generated by nodes of the hidden layer of nodes;
selecting a subset of training samples based on distances between vectors representing the training samples;
for each selected training sample, associating the selected training sample with a label representing an expected result that should be output by the neural network for the selected training sample; and
providing the selected training samples of the subset as a second training dataset for retraining the neural network.
12 . The non-transitory computer readable medium of claim 1 , wherein selecting the subset of training samples comprises:
clustering the training samples to generate a plurality of clusters, the clustering associating each training sample with a cluster, the clustering based on distances between the vectors representing the training samples; for each training sample, determining a score indicative of similarity of the training sample with the cluster associated with the training sample; and excluding one or more training samples based on the scores to determine the subset of training samples.
13 . The non-transitory computer readable medium of claim 12 , wherein the one or more training samples are excluded responsive to each of the one or more training samples having a score indicating high similarity of the training sample to the cluster associated with the training sample.
14 . The non-transitory computer readable medium of claim 1 , wherein the plurality of layers of nodes represents a sequence of layers of nodes comprising a plurality of hidden layers, wherein the hidden layer of claim 1 is the last hidden layer from the plurality of hidden layers.
15 . The non-transitory computer readable medium of claim 1 , wherein the neural network further comprises an output layer of nodes for outputting a result of executing the neural network, wherein the hidden layer represents nodes that provide inputs to the output layer of nodes.
16 . The non-transitory computer readable medium of claim 1 , wherein the training process further comprises:
receiving a second plurality of training samples; executing the neural network to process the second plurality of training samples; for each of the training sample in the second plurality of training samples, receiving a second vector representing the training sample, the second vector generated by the nodes of the hidden layer of nodes; selecting a second subset of training samples from the second plurality of training samples based on distances between the vectors representing the training samples in the second plurality of training samples; and providing the second subset as a third training dataset for retraining the neural network.
17 . The non-transitory computer readable medium of claim 16 , wherein the training process is repeatedly performed until an aggregate measure of difference between a result outputted by the neural network and the expected result is below a threshold value.
18 . The non-transitory computer readable medium of claim 1 , wherein the neural network represents a deep learning model.
19 . The non-transitory computer readable medium of claim 1 , wherein the target sample is a natural language description of an input image.
20 . The non-transitory computer readable medium of claim 1 , wherein the neural network is an autoencoder configured to receive an input and regenerate the input as output of the neural network.Join the waitlist — get patent alerts
Track US2021117802A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.