US2019126472A1PendingUtilityA1

Reinforcement and imitation learning for a task

Assignee: DEEPMIND TECH LTDPriority: Oct 27, 2017Filed: Oct 29, 2018Published: May 2, 2019
Est. expiryOct 27, 2037(~11.2 yrs left)· nominal 20-yr term from priority
G06N 3/045G06N 3/044G06N 3/092G06N 3/0442G06N 3/08B25J 9/163B25J 9/1697B25J 9/161G06N 3/098G06N 3/094G06N 3/09G06N 3/0464G06N 3/008G06N 3/084
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A neural network control system for controlling an agent to perform a task in a real-world environment, operates based on both image data and proprioceptive data describing the configuration of the agent. The training of the control system includes both imitation learning, using datasets generated from previous performances of the task, and reinforcement learning, based on rewards calculated from control data output by the control system.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method of training a neural network to generate commands for controlling an agent to perform a task in an environment, the method comprising:
 obtaining, for each of a plurality of performances of the task, a respective dataset characterizing the corresponding performance of the task; and   using the dataset, training a neural network to generate commands for controlling the agent based on image data encoding captured images of the environment and proprioceptive data comprising one or more variables describing configurations of the agent;   wherein the training the neural network comprises:   using the neural network to generate a plurality of sets of one or more commands,   for each set of commands generating at least one corresponding reward value indicative of how successfully the task is carried out upon implementation of the set of commands by the agent, and   adjusting one or more parameters of the neural network based on the datasets, the sets of commands and the corresponding reward values.   
     
     
         2 . The method of  claim 1  in which adjusting the one or more parameters of the neural network comprises adjusting the neural network based on a hybrid energy function, the hybrid energy function including both an imitation reward value derived using the datasets and the generated sets of commands, and task reward term calculated using the generated reward values. 
     
     
         3 . The method of the  claim 2  including using the datasets to generate a discriminator network, and deriving the imitation reward value using the discriminator network and the sets of one or more commands. 
     
     
         4 . The method of  claim 3  in which the discriminator network receives data characterizing the positions of objects in the environment. 
     
     
         5 . The method of  claim 1 , in which the reward value is generated by computationally simulating a process carried out by the agent in the environment based on the corresponding set of commands to generate a final state of the environment, and calculating an initial reward value based at least on the final state of the environment. 
     
     
         6 . The method of  claim 5 , in which updates to the neural network are calculated using an activation function estimator obtained by subtracting a value function from the initial reward value, and the initial reward value is calculated according to a task reward function based on the final state of the environment. 
     
     
         7 . The method of  claim 6  in which the value function is calculated using data characterizing the positions of objects in the environment. 
     
     
         8 . The method of  claim 6  in which the value function is calculated by an adaptive model. 
     
     
         9 . The method of  claim 1 , in which the neural network comprises a convolutional neural network which receives the image data and from it generates convolved data, the neural network further comprising at least one adaptive component which receives the output of the convolutional neural network and the proprioceptive data. 
     
     
         10 . The method according to  claim 9  in which the adaptive component is a perceptron. 
     
     
         11 . The method of  claim 9  in which the neural network further comprises a recursive neural network, which receives input data generated both from the image data and the proprioceptive data. 
     
     
         12 . The method of  claim 9 , further including defining at least one auxiliary task, and training the convolutional network as part of an adaptive system which is trained to perform the auxiliary task based on image data. 
     
     
         13 . The method of  claim 1 , in which the training of the neural network is performed in parallel with the training of a plurality of additional instances of the neural network by respective workers, the adjustment of the parameters of the neural network being additionally based on reward values indicative of how successfully the task is carried out by simulated agents based on sets of commands generated by the additional neural networks. 
     
     
         14 . The method of  claim 1 , in which the step of using the neural network to generate a plurality of sets of commands is performed at least once by supplying to the neural network image data and proprioceptive data which characterizes a state associated with one of the performances of the task. 
     
     
         15 . The method of  claim 1 , further comprising, prior to training the neural network, defining a plurality of stages of the task, and for each stage of the task defining a respective plurality of initial states,
 the step of using the neural network to generate a plurality of sets of commands being performed at least once, for each task stage, by supplying to the neural network image data and proprioceptive data which characterizes one of the corresponding plurality of initial states.   
     
     
         16 . A method of performing a task, the method comprising:
 training a neural network to generate commands for controlling an agent to perform the task in an environment, by a method according to any preceding claim; and   a plurality of times performing the steps of:   (i) capturing images of an environment and generating image data encoding the images;   (ii) capturing proprioceptive data comprising one or more variables describing configurations of the agent;   (iii) transmitting the image data and the proprioceptive data to the neural network, the neural network generating at least one command based on the image data and the proprioceptive data; and   (iv) transmitting the command to the agent, the agent being operative to perform the command within the environment;   whereby the neural network successively generates a sequence of commands to control the agent to perform the task.   
     
     
         17 . The method of  claim 16  in which the step of obtaining, for each of a plurality of performances of the task, a respective dataset characterizing the corresponding performance of the task, is performed by controlling the agent to perform the task a plurality of times, and for each performance generating a respective dataset characterizing the performance. 
     
     
         18 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:
 obtaining, for each of a plurality of performances of the task, a respective dataset characterizing the corresponding performance of the task; and   using the dataset, training a neural network to generate commands for controlling the agent based on image data encoding captured images of the environment and proprioceptive data comprising one or more variables describing configurations of the agent   wherein the training the neural network comprises:   using the neural network to generate a plurality of sets of one or more commands,   for each set of commands generating at least one corresponding reward value indicative of how successfully the task is carried out upon implementation of the set of commands by the agent, and   adjusting one or more parameters of the neural network based on the datasets, the sets of commands and the corresponding reward values.   
     
     
         19 . The system of  claim 18  further including: an agent operative to perform commands generated by the neural network; at least one image capture device operative to capture images of an environment and generate image data encoding the images; and at least one device operative to capture proprioceptive data comprising the one or more variables describing configurations of the agent. 
     
     
         20 . (canceled)

Join the waitlist — get patent alerts

Track US2019126472A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.