Framework for Learning to Transfer Learn
Abstract
A method includes receiving a source data set and a target data set and identifying a loss function for a deep learning model based on the source data set and the target data set. The loss function includes encoder weights, source classifier layer weights, target classifier layer weights, coefficients, and a policy weight. During a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL) architecture, the method also includes: applying gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier weights that minimize the loss function; and determining the coefficients by sampling actions of a policy model. During a second phase of each of the plurality of learning iterations, determining the policy weight that maximizes an evaluation metric.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving, at data processing hardware, a source data set and a target data set; identifying, by the data processing hardware, a loss function for a deep learning model based on the source data set and the target data set, the loss function comprising:
encoder weights;
source classifier layer weights;
target classifier layer weights;
coefficients; and
a policy weight,
during a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL ) architecture configured to learn weight assignments for the deep learning model:
applying, by the data processing hardware, gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function; and
determining, by the data processing hardware, the coefficients by sampling actions of a policy model; and
during a second phase of each of the plurality of learning iterations for the L2TL architecture, determining, by the data processing hardware, the policy weight that maximizes an evaluation metric for the loss function.
2 . The method of claim 1 , wherein the policy model is fixed while performing the first phase of the learning iteration.
3 . The method of claim 1 , wherein the policy model comprises a reinforcement learning-based policy model.
4 . The method of claim 1 , wherein determining the policy weight that maximizes the evaluation metric for the loss function comprises using the encoder weights and the target classification layer weights learned during the first phase.
5 . The method of claim 1 , wherein the evaluation metric for the loss function quantifies performance of the deep learning model on a target evaluation dataset, the target evaluation dataset comprising a subset of data samples in the target dataset that were not previously seen by the deep learning model.
6 . The method of claim 1 , further comprising, during the first phase of each of the plurality of learning iterations:
sampling, by the data processing hardware, a training batch of source data samples from the source data set having a particular size; and selecting, by the data processing hardware, the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function.
7 . The method of claim 1 , further comprising, during the second phase of each of the plurality of learning iterations:
training, by the data processing hardware, the policy model using policy gradient on a target evaluation dataset to compute a reward that maximizes the evaluation metric, wherein determining the policy weight that maximizes the evaluation metric for the loss function is based on the computed reward.
8 . The method of claim 1 , wherein:
the source data set comprises a first plurality of images; and the target data set comprises a second plurality of images.
9 . The method of claim 8 , wherein a number of images in the first plurality of images of the source data set is greater than a number of images in the second plurality of images of the target data set.
10 . The method of claim 1 , wherein the L2TL architecture comprises an encoder network layer, a source classifier layer, and a target classifier layer.
11 . A system comprising:
data processing hardware, and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving a source data set and a target data set;
identifying a loss function for a deep learning model based on the source data set and the target data set, the loss function comprising:
encoder weights;
source classifier layer weights;
target classifier layer weights;
coefficients; and
a policy weight;
during a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL) architecture configured to learn weight assignments for the deep learning model:
applying gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function; and
determining the coefficients by sampling actions of a policy model; and
during a second phase of each of the plurality of learning iterations for the L2TL architecture, determining the policy weight that maximizes an evaluation metric for the loss function.
12 . The system of claim 11 , wherein the policy model is fixed while performing the first phase of the learning iteration.
13 . The system of claim 11 , wherein the policy model comprises a reinforcement learning-based policy model.
14 . The system of claim 11 , wherein determining the policy weight that maximizes the evaluation metric for the loss function comprises using the encoder weights learned during the first phase.
15 . The system of claim 11 , wherein the evaluation metric for the loss function quantifies performance of the deep learning model on a target evaluation dataset, the target evaluation dataset comprising a subset of data samples in the target dataset that were not previously seen by the deep learning model.
16 . The system of claim 11 , wherein the operations further comprise, during the first phase of each of the plurality of learning iterations:
sampling a training batch of source data samples from the source data set having a particular size, and selecting the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function.
17 . The system of claim 11 , wherein the operations further comprise, during the second phase of each of the plurality of learning iterations:
training the policy model using policy gradient on a target evaluation dataset to compute a reward that maximizes the evaluation metric, wherein determining the policy weight that maximizes the evaluation metric for the loss function is based on the computed reward.
18 . The system of claim 11 , wherein:
the source data set comprises a first plurality of images; and the target data set comprises a second plurality of images.
19 . The system of claim 18 , wherein a number of images in the first plurality of images of the source data set is greater than a number of images in the second plurality of images of the target data set.
20 . The system of claim 11 , wherein the L2TL architecture comprises an encoder network layer, a source classifier layer, and a target classifier layer.Join the waitlist — get patent alerts
Track US2021034976A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.