Training neural networks using learned adaptive learning rates
Abstract
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network. One of the methods includes training the neural network for one or more training steps in accordance with a current learning rate; generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps; providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate; obtaining as output from the controller neural network the controller output that defines the updated learning rate; and setting the learning rate to the updated learning rate.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of training a trainee neural network having a plurality of trainee parameters by repeatedly adjusting values of the trainee network parameters, the method comprising repeatedly performing operations comprising:
training the trainee neural network for one or more training steps, the training comprising, at each training step:
receiving a plurality of training inputs;
determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters;
applying a learning rate to the gradient to generate a trainee parameter value update; and
updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters;
generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps; providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate; obtaining as output from the controller neural network the controller output that defines the updated learning rate; and setting the learning rate to the updated learning rate.
2 . The method of claim 1 , wherein the controller output is a scaling factor to be applied to the learning rate used for the one or more training steps to generate the updated learning rate.
3 . The method of claim 1 , wherein training dynamics observation comprises the learning rate used for the one or more training steps.
4 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on a current training loss of the trainee neural network on the training inputs for the one or more training steps.
5 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on a current validation loss of the trainee neural network on validation data.
6 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the updated values of the parameters of a designated layer in the trainee neural network.
7 . The method of claim 1 , wherein the training dynamics observation comprises one or more features that are each based on outputs generated by the trainee neural network for the training inputs for the one or more training steps.
8 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the training inputs for the one or more training steps.
9 . The method of claim 1 , wherein the controller neural network has been trained jointly with the training of a second, different neural network that has a different architecture from the trainee neural network and the values of the parameters of the controller neural network are fixed during the training of the trainee neural network.
10 . The method of claim 1 , the operations further comprising:
obtaining one or more rewards that measure the performance of the trainee neural network after the one or more training steps; and updating the values of the parameters of the controller neural network by training, based on the reward, the controller neural network through reinforcement learning to maximize an objective function that measures the expected time discounted reward during the training of the trainee neural network.
11 . The method of claim 10 , wherein the reward is based on a validation loss of the trainee neural network after the training for the one or more training steps.
12 . The method of claim 10 , wherein training the controller neural network comprises training the trainee neural network using a policy gradient reinforcement learning technique.
13 . The method of claim 12 , wherein the policy gradient reinforcement learning technique is proximal policy optimization (PPO).
14 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to train a trainee neural network having a plurality of trainee parameters by repeatedly performing operations to adjust values of the trainee network parameters, the operations comprising:
training the trainee neural network for one or more training steps, the training comprising, at each training step:
receiving a plurality of training inputs;
determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters;
applying a learning rate to the gradient to generate a trainee parameter value update; and
updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters;
generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps; providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate; obtaining as output from the controller neural network the controller output that defines the updated learning rate; and setting the learning rate to the updated learning rate.
15 . The method of claim 1 , wherein the controller output is a scaling factor to be applied to the learning rate used for the one or more training steps to generate the updated learning rate.
16 . The method of claim 1 , wherein training dynamics observation comprises the learning rate used for the one or more training steps.
17 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on a current training loss of the trainee neural network on the training inputs for the one or more training steps.
18 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on a current validation loss of the trainee neural network on validation data.
19 . The method of claim 1 , wherein the training dynamics observation comprises a feature that is based on statistics of the updated values of the parameters of a designated layer in the trainee neural network.
20 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to train a trainee neural network having a plurality of trainee parameters by repeatedly performing operations to adjust values of the trainee network parameters, the operations comprising:
training the trainee neural network for one or more training steps, the training comprising, at each training step:
receiving a plurality of training inputs;
determining, based on processing the training inputs using the trainee neural network and in accordance with current values of the trainee parameters, a gradient of an objective function with respect to the trainee network parameters;
applying a learning rate to the gradient to generate a trainee parameter value update; and
updating the current values of the trainee network parameters by applying the trainee parameter value update to the current values of the trainee parameters;
generating a training dynamics observation characterizing the training of the trainee neural network on the one or more training steps; providing the training dynamics observation as input to a controller neural network that is configured to process the training dynamics observation to generate a controller output that defines an updated learning rate; obtaining as output from the controller neural network the controller output that defines the updated learning rate; and setting the learning rate to the updated learning rate.Join the waitlist — get patent alerts
Track US2021034973A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.