US2021034976A1PendingUtilityA1

Framework for Learning to Transfer Learn

Assignee: GOOGLE LLCPriority: Aug 2, 2019Filed: Aug 2, 2020Published: Feb 4, 2021
Est. expiryAug 2, 2039(~13 yrs left)· nominal 20-yr term from priority
G06F 18/241G06N 3/045G06N 3/096G06N 3/0985G06N 3/092G06N 3/09G06N 20/00G06N 3/08G06N 3/04
52
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes receiving a source data set and a target data set and identifying a loss function for a deep learning model based on the source data set and the target data set. The loss function includes encoder weights, source classifier layer weights, target classifier layer weights, coefficients, and a policy weight. During a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL) architecture, the method also includes: applying gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier weights that minimize the loss function; and determining the coefficients by sampling actions of a policy model. During a second phase of each of the plurality of learning iterations, determining the policy weight that maximizes an evaluation metric.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 receiving, at data processing hardware, a source data set and a target data set;   identifying, by the data processing hardware, a loss function for a deep learning model based on the source data set and the target data set, the loss function comprising:
 encoder weights; 
 source classifier layer weights; 
 target classifier layer weights; 
 coefficients; and 
 a policy weight, 
   during a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL ) architecture configured to learn weight assignments for the deep learning model:
 applying, by the data processing hardware, gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function; and 
 determining, by the data processing hardware, the coefficients by sampling actions of a policy model; and 
   during a second phase of each of the plurality of learning iterations for the L2TL architecture, determining, by the data processing hardware, the policy weight that maximizes an evaluation metric for the loss function.   
     
     
         2 . The method of  claim 1 , wherein the policy model is fixed while performing the first phase of the learning iteration. 
     
     
         3 . The method of  claim 1 , wherein the policy model comprises a reinforcement learning-based policy model. 
     
     
         4 . The method of  claim 1 , wherein determining the policy weight that maximizes the evaluation metric for the loss function comprises using the encoder weights and the target classification layer weights learned during the first phase. 
     
     
         5 . The method of  claim 1 , wherein the evaluation metric for the loss function quantifies performance of the deep learning model on a target evaluation dataset, the target evaluation dataset comprising a subset of data samples in the target dataset that were not previously seen by the deep learning model. 
     
     
         6 . The method of  claim 1 , further comprising, during the first phase of each of the plurality of learning iterations:
 sampling, by the data processing hardware, a training batch of source data samples from the source data set having a particular size; and   selecting, by the data processing hardware, the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function.   
     
     
         7 . The method of  claim 1 , further comprising, during the second phase of each of the plurality of learning iterations:
 training, by the data processing hardware, the policy model using policy gradient on a target evaluation dataset to compute a reward that maximizes the evaluation metric,   wherein determining the policy weight that maximizes the evaluation metric for the loss function is based on the computed reward.   
     
     
         8 . The method of  claim 1 , wherein:
 the source data set comprises a first plurality of images; and   the target data set comprises a second plurality of images.   
     
     
         9 . The method of  claim 8 , wherein a number of images in the first plurality of images of the source data set is greater than a number of images in the second plurality of images of the target data set. 
     
     
         10 . The method of  claim 1 , wherein the L2TL architecture comprises an encoder network layer, a source classifier layer, and a target classifier layer. 
     
     
         11 . A system comprising:
 data processing hardware, and   memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
 receiving a source data set and a target data set; 
 identifying a loss function for a deep learning model based on the source data set and the target data set, the loss function comprising:
 encoder weights; 
 source classifier layer weights; 
 target classifier layer weights; 
 coefficients; and 
 a policy weight; 
 
 during a first phase of each of a plurality of learning iterations for a learning to transfer learn (L2TL) architecture configured to learn weight assignments for the deep learning model:
 applying gradient decent-based optimization to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function; and 
 determining the coefficients by sampling actions of a policy model; and 
 
 during a second phase of each of the plurality of learning iterations for the L2TL architecture, determining the policy weight that maximizes an evaluation metric for the loss function. 
   
     
     
         12 . The system of  claim 11 , wherein the policy model is fixed while performing the first phase of the learning iteration. 
     
     
         13 . The system of  claim 11 , wherein the policy model comprises a reinforcement learning-based policy model. 
     
     
         14 . The system of  claim 11 , wherein determining the policy weight that maximizes the evaluation metric for the loss function comprises using the encoder weights learned during the first phase. 
     
     
         15 . The system of  claim 11 , wherein the evaluation metric for the loss function quantifies performance of the deep learning model on a target evaluation dataset, the target evaluation dataset comprising a subset of data samples in the target dataset that were not previously seen by the deep learning model. 
     
     
         16 . The system of  claim 11 , wherein the operations further comprise, during the first phase of each of the plurality of learning iterations:
 sampling a training batch of source data samples from the source data set having a particular size, and   selecting the source data samples from the training batch of source data samples that have the N-best confidence scores for use in training the deep learning model to learn the encoder weights, the source classifier layer weights, and the target classifier layer weights that minimize the loss function.   
     
     
         17 . The system of  claim 11 , wherein the operations further comprise, during the second phase of each of the plurality of learning iterations:
 training the policy model using policy gradient on a target evaluation dataset to compute a reward that maximizes the evaluation metric,   wherein determining the policy weight that maximizes the evaluation metric for the loss function is based on the computed reward.   
     
     
         18 . The system of  claim 11 , wherein:
 the source data set comprises a first plurality of images; and   the target data set comprises a second plurality of images.   
     
     
         19 . The system of  claim 18 , wherein a number of images in the first plurality of images of the source data set is greater than a number of images in the second plurality of images of the target data set. 
     
     
         20 . The system of  claim 11 , wherein the L2TL architecture comprises an encoder network layer, a source classifier layer, and a target classifier layer.

Join the waitlist — get patent alerts

Track US2021034976A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.