US2017032245A1PendingUtilityA1
Systems and Methods for Providing Reinforcement Learning in a Deep Learning System
Est. expiryJul 1, 2035(~8.9 yrs left)· nominal 20-yr term from priority
G06N 3/08G06N 3/044G06N 3/045G06N 3/092G06N 3/0464G06N 99/005
39
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods for providing reinforcement learning for a deep learning network are disclosed. A reinforcement learning process that provides deep exploration is provided by a bootstrap that applied to a sample of observed and artificial data to facilitate deep exploration via a Thompson sampling approach.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A deep learning system comprising:
at least one processor; memory accessible by each at least one processor; instructions that when read by the at least one processor direct the at least one processor to:
maintain a deep neural network; and
apply a reinforcement learning process to the deep neural network where the reinforcement learning process includes:
receive a set of observed data and a set artificial data,
for each of one or more episodes:
sample from a set of data that is a union of the set of observed data and the set of artificial data to generate set of training data;
determine a state-action value function for the set of training data using a bootstrap process and an approximator where the approximator that estimates a state-action function for a dataset;
for each time step in each one or more episode:
determine a state of the system for a current time step from the set of training data;
select an action based on the determined state of the system and a policy mapping actions to the state of the system;
determine results for the action including a reward and a transition state that result from the selected action; and
store result data for the current time step that includes the state, the action, the transition state, and
update the set of the observed data with the result data from at least one time step for each of the one or more the episodes.
2 . The deep learning system of claim 1 wherein the instructions further direct the at least one processor to generate the set of artificial data from the set of observed data.
3 . The deep learning system of claim 2 wherein the instructions to generate the artificial data include instruction that direct the at least one processor:
sample the set of observed data with replacement to generate the set of artificial data.
4 . The deep learning system of claim 2 wherein the instructions to generate the artificial data include instructions that direct the at least one processor to:
sample a plurality state-action pairs from a diffusely mixed generative model; and
assign each of the plurality of sampled state-action pairs stochastically optimistic rewards and random state transitions.
5 . The deep learning system of claim 1 wherein the instructions further direct the at least one processor to:
maintain a training mask that indicates the result data from each of the time period in each episode to be used in training; and
wherein the updating of the set of observed data includes adding the result data from each time period of an episode indicated in the training mask.
6 . The deep learning network of claim 1 where the instructions further direct the processor to:
receive the approximator as an input.
7 . The deep learning network of claim 1 wherein the instructions further direct the processor to:
read the approximator from memory.
8 . The deep learning network of claim 1 wherein the approximator is a neural network trained to fit a state-action value function to the data set via a least squared iteration.
9 . The deep learning network of claim 1 wherein a plurality of reinforcement learning processes are applied to the deep neural network.
10 . The deep learning network of claim 9 wherein each of the plurality of reinforcement learning processes independently maintain the set of observed data.
11 . The deep learning network of claim 9 wherein the plurality of reinforcement learning processes cooperatively maintain the set of observed data.
12 . The deep learning process of claim 9 wherein the instruction further direct the processor to:
maintain a bootstrap mask that indicates each element in the set of observed data that is available to each of the plurality of reinforcement learning process.
13 . A method performed by at least one processor executing instructions stored in memory to perform the method to provide reinforcement learning in a deep learning network, the method comprising:
receiving a set of observed data and a set artificial data;
for each of one or more episodes:
sampling from a set of data that is a union of the set of observed data and the set of artificial data to generate set of training data,
determining a state-action value function for the set of training data using a bootstrap process and an approximator where the approximator that estimates a state-action function for a dataset,
for each time step in each one or more episode:
determining a state of the system for a current time step from the set of training data;
selecting an action based on the determined state of the system and a policy mapping actions to the state of the system;
determining results for the action including a reward and a transition state that result from the selected action; and
storing result data for the current time step that includes the state, the action, the transition state, and
updating the set of the observed data with the result data from at least one time step of each of the one or more episodes.
14 . The method of claim 13 further comprising generating the set of artificial data from the set of observed data.
15 . The method of claim 14 further comprising:
sampling the set of observed data with replacement to generate the set of artificial data.
16 . The method of claim 14 further comprising:
sampling a plurality state-action pairs from a diffusely mixed generative model; and
assigning each of the plurality of sampled state-action pairs stochastically optimistic rewards and random state transitions.
17 . The method of claim 13 further comprising:
maintaining a training mask that indicates the result data from each of the time period in each episode to be used in training; and
wherein the updating of the set of observed data includes adding the result data from each time period of an episode indicated in the training mask.
18 . The method of claim 13 further comprising:
receiving the approximator as an input.
19 . The method of claim 13 further comprising:
read the approximator from memory.
20 . The method of claim 13 wherein the approximator is a neural network trained to fit a state-action value function to the data set via a least squared iteration.
21 . The method of claim 13 wherein a plurality of reinforcement learning methods are applied to the deep neural network.
22 . The method of claim 21 wherein each of the plurality of reinforcement learning methods independently maintain the set of observed data.
23 . The method of claim 21 wherein the plurality of reinforcement learning methods cooperatively maintain the set of observed data.
24 . The method of claim 21 further comprising:
maintaining a bootstrap mask that indicates each element in the set of observed data that is available to each of the plurality of reinforcement learning process.Join the waitlist — get patent alerts
Track US2017032245A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.