US2017032245A1PendingUtilityA1

Systems and Methods for Providing Reinforcement Learning in a Deep Learning System

Assignee: UNIV LELAND STANFORD JUNIORPriority: Jul 1, 2015Filed: Jul 15, 2016Published: Feb 2, 2017
Est. expiryJul 1, 2035(~8.9 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/044G06N 3/045G06N 3/092G06N 3/0464G06N 99/005
39
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for providing reinforcement learning for a deep learning network are disclosed. A reinforcement learning process that provides deep exploration is provided by a bootstrap that applied to a sample of observed and artificial data to facilitate deep exploration via a Thompson sampling approach.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A deep learning system comprising:
 at least one processor;   memory accessible by each at least one processor;   instructions that when read by the at least one processor direct the at least one processor to:
 maintain a deep neural network; and 
 apply a reinforcement learning process to the deep neural network where the reinforcement learning process includes:
 receive a set of observed data and a set artificial data, 
 for each of one or more episodes:
 sample from a set of data that is a union of the set of observed data and the set of artificial data to generate set of training data; 
 
 determine a state-action value function for the set of training data using a bootstrap process and an approximator where the approximator that estimates a state-action function for a dataset; 
 for each time step in each one or more episode:
 determine a state of the system for a current time step from the set of training data; 
 select an action based on the determined state of the system and a policy mapping actions to the state of the system; 
 determine results for the action including a reward and a transition state that result from the selected action; and 
 store result data for the current time step that includes the state, the action, the transition state, and 
 
 update the set of the observed data with the result data from at least one time step for each of the one or more the episodes. 
 
   
     
     
         2 . The deep learning system of  claim 1  wherein the instructions further direct the at least one processor to generate the set of artificial data from the set of observed data. 
     
     
         3 . The deep learning system of  claim 2  wherein the instructions to generate the artificial data include instruction that direct the at least one processor:
 sample the set of observed data with replacement to generate the set of artificial data. 
 
     
     
         4 . The deep learning system of  claim 2  wherein the instructions to generate the artificial data include instructions that direct the at least one processor to:
 sample a plurality state-action pairs from a diffusely mixed generative model; and 
 assign each of the plurality of sampled state-action pairs stochastically optimistic rewards and random state transitions. 
 
     
     
         5 . The deep learning system of  claim 1  wherein the instructions further direct the at least one processor to:
 maintain a training mask that indicates the result data from each of the time period in each episode to be used in training; and 
 wherein the updating of the set of observed data includes adding the result data from each time period of an episode indicated in the training mask. 
 
     
     
         6 . The deep learning network of  claim 1  where the instructions further direct the processor to:
 receive the approximator as an input. 
 
     
     
         7 . The deep learning network of  claim 1  wherein the instructions further direct the processor to:
 read the approximator from memory. 
 
     
     
         8 . The deep learning network of  claim 1  wherein the approximator is a neural network trained to fit a state-action value function to the data set via a least squared iteration. 
     
     
         9 . The deep learning network of  claim 1  wherein a plurality of reinforcement learning processes are applied to the deep neural network. 
     
     
         10 . The deep learning network of  claim 9  wherein each of the plurality of reinforcement learning processes independently maintain the set of observed data. 
     
     
         11 . The deep learning network of  claim 9  wherein the plurality of reinforcement learning processes cooperatively maintain the set of observed data. 
     
     
         12 . The deep learning process of  claim 9  wherein the instruction further direct the processor to:
 maintain a bootstrap mask that indicates each element in the set of observed data that is available to each of the plurality of reinforcement learning process. 
 
     
     
         13 . A method performed by at least one processor executing instructions stored in memory to perform the method to provide reinforcement learning in a deep learning network, the method comprising:
 receiving a set of observed data and a set artificial data;
 for each of one or more episodes:
 sampling from a set of data that is a union of the set of observed data and the set of artificial data to generate set of training data, 
 determining a state-action value function for the set of training data using a bootstrap process and an approximator where the approximator that estimates a state-action function for a dataset, 
 
 for each time step in each one or more episode:
 determining a state of the system for a current time step from the set of training data; 
 selecting an action based on the determined state of the system and a policy mapping actions to the state of the system; 
 determining results for the action including a reward and a transition state that result from the selected action; and 
 storing result data for the current time step that includes the state, the action, the transition state, and 
 
 updating the set of the observed data with the result data from at least one time step of each of the one or more episodes. 
   
     
     
         14 . The method of  claim 13  further comprising generating the set of artificial data from the set of observed data. 
     
     
         15 . The method of  claim 14  further comprising:
 sampling the set of observed data with replacement to generate the set of artificial data. 
 
     
     
         16 . The method of  claim 14  further comprising:
 sampling a plurality state-action pairs from a diffusely mixed generative model; and 
 assigning each of the plurality of sampled state-action pairs stochastically optimistic rewards and random state transitions. 
 
     
     
         17 . The method of  claim 13  further comprising:
 maintaining a training mask that indicates the result data from each of the time period in each episode to be used in training; and 
 wherein the updating of the set of observed data includes adding the result data from each time period of an episode indicated in the training mask. 
 
     
     
         18 . The method of  claim 13  further comprising:
 receiving the approximator as an input. 
 
     
     
         19 . The method of  claim 13  further comprising:
 read the approximator from memory. 
 
     
     
         20 . The method of  claim 13  wherein the approximator is a neural network trained to fit a state-action value function to the data set via a least squared iteration. 
     
     
         21 . The method of  claim 13  wherein a plurality of reinforcement learning methods are applied to the deep neural network. 
     
     
         22 . The method of  claim 21  wherein each of the plurality of reinforcement learning methods independently maintain the set of observed data. 
     
     
         23 . The method of  claim 21  wherein the plurality of reinforcement learning methods cooperatively maintain the set of observed data. 
     
     
         24 . The method of  claim 21  further comprising:
 maintaining a bootstrap mask that indicates each element in the set of observed data that is available to each of the plurality of reinforcement learning process.

Join the waitlist — get patent alerts

Track US2017032245A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.