US2021034973A1PendingUtilityA1

Training neural networks using learned adaptive learning rates

Assignee: GOOGLE LLCPriority: Jul 30, 2019Filed: Jul 30, 2020Published: Feb 4, 2021
Est. expiryJul 30, 2039(~13 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/044G06N 3/08G06N 3/0442G06N 3/0985G06N 3/096G06N 3/09G06N 3/0895G06N 3/0464G06N 3/092G06N 3/0454
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network. One of the methods includes training the neural network for one or more training steps in accordance with a current learning rate; generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps; providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate; obtaining as output from the controller neural network the controller output that defines the updated learning rate; and setting the learning rate to the updated learning rate.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of training a trainee neural network having a plurality of trainee parameters by repeatedly adjusting values of the trainee network parameters, the method comprising repeatedly performing operations comprising:
 training the trainee neural network for one or more training steps, the training comprising, at each training step:
 receiving a plurality of training inputs; 
 determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters; 
 applying a learning rate to the gradient to generate a trainee parameter value update; and 
 updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters; 
   generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps;   providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate;   obtaining as output from the controller neural network the controller output that defines the updated learning rate; and   setting the learning rate to the updated learning rate.   
     
     
         2 . The method of  claim 1 , wherein the controller output is a scaling factor to be applied to the learning rate used for the one or more training steps to generate the updated learning rate. 
     
     
         3 . The method of  claim 1 , wherein training dynamics observation comprises the learning rate used for the one or more training steps. 
     
     
         4 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on a current training loss of the trainee neural network on the training inputs for the one or more training steps. 
     
     
         5 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on a current validation loss of the trainee neural network on validation data. 
     
     
         6 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the updated values of the parameters of a designated layer in the trainee neural network. 
     
     
         7 . The method of  claim 1 , wherein the training dynamics observation comprises one or more features that are each based on outputs generated by the trainee neural network for the training inputs for the one or more training steps. 
     
     
         8 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the training inputs for the one or more training steps. 
     
     
         9 . The method of  claim 1 , wherein the controller neural network has been trained jointly with the training of a second, different neural network that has a different architecture from the trainee neural network and the values of the parameters of the controller neural network are fixed during the training of the trainee neural network. 
     
     
         10 . The method of  claim 1 , the operations further comprising:
 obtaining one or more rewards that measure the performance of the trainee neural network after the one or more training steps; and   updating the values of the parameters of the controller neural network by training, based on the reward, the controller neural network through reinforcement learning to maximize an objective function that measures the expected time discounted reward during the training of the trainee neural network.   
     
     
         11 . The method of  claim 10 , wherein the reward is based on a validation loss of the trainee neural network after the training for the one or more training steps. 
     
     
         12 . The method of  claim 10 , wherein training the controller neural network comprises training the trainee neural network using a policy gradient reinforcement learning technique. 
     
     
         13 . The method of  claim 12 , wherein the policy gradient reinforcement learning technique is proximal policy optimization (PPO). 
     
     
         14 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to train a trainee neural network having a plurality of trainee parameters by repeatedly performing operations to adjust values of the trainee network parameters, the operations comprising:
 training the trainee neural network for one or more training steps, the training comprising, at each training step:
 receiving a plurality of training inputs; 
 determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters; 
 applying a learning rate to the gradient to generate a trainee parameter value update; and 
 updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters; 
   generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps;   providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate;   obtaining as output from the controller neural network the controller output that defines the updated learning rate; and   setting the learning rate to the updated learning rate.   
     
     
         15 . The method of  claim 1 , wherein the controller output is a scaling factor to be applied to the learning rate used for the one or more training steps to generate the updated learning rate. 
     
     
         16 . The method of  claim 1 , wherein training dynamics observation comprises the learning rate used for the one or more training steps. 
     
     
         17 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on a current training loss of the trainee neural network on the training inputs for the one or more training steps. 
     
     
         18 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on a current validation loss of the trainee neural network on validation data. 
     
     
         19 . The method of  claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the updated values of the parameters of a designated layer in the trainee neural network. 
     
     
         20 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to train a trainee neural network having a plurality of trainee parameters by repeatedly performing operations to adjust values of the trainee network parameters, the operations comprising:
 training the trainee neural network for one or more training steps, the training comprising, at each training step:
 receiving a plurality of training inputs; 
 determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters; 
 applying a learning rate to the gradient to generate a trainee parameter value update; and 
 updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters; 
   generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps;   providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate;   obtaining as output from the controller neural network the controller output that defines the updated learning rate; and   setting the learning rate to the updated learning rate.

Join the waitlist — get patent alerts

Track US2021034973A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.