Framework and methods of diverse exploration for fast and safe policy improvement
Abstract
The present technology addresses the problem of quickly and safely improving policies in online reinforcement learning domains. As its solution, an exploration strategy comprising diverse exploration (DE) is employed, which learns and deploys a diverse set of safe policies to explore the environment. DE theory explains why diversity in behavior policies enables effective exploration without sacrificing exploitation. An empirical study shows that an online policy improvement algorithm framework implementing the DE strategy can achieve both fast policy improvement and safe online performance.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for learning and deploying a set of behavior policies for an artificial agent, selected from a space of behavior policies, each respective behavior policy having a statistically expected return no worse than a lower bound of policy performance which excludes a portion of the set of behavior policies, comprising:
at least one automated processor configured to:
iteratively improve a behavior policy, wherein each iteration:
employs a diverse exploration strategy which strives for behavior diversity in a space of stochastic policies by deploying a diverse set comprising a plurality of behavior policies which are ensured as being safe during each iteration of policy improvement; and
assesses policy performance of each behavior policy of the diverse set with respect to the artificial agent for each iteration of the iterative improvement.
2 . The system according to claim 1 , wherein each behavior policy has a variance associated with the estimate of its policy performance by importance sampling, and each diverse set has a common average variance, in each of a plurality of policy improvement iterations.
3 . The system according to claim 1 , wherein each of the policy performance and policy behavior diversity is quantified according to a common objective function, and the at least one automated processor is further configured to employ the common objective function to assess the policy performance of the diverse set.
4 . The system according to claim 1 , wherein the diverse set comprises a plurality of behavior policies predefined upon commencement of a respective single iteration of policy improvement.
5 . The system according to claim 1 , wherein the at least one processor is further configured to adaptively define the diverse set based on assessed policy performance within a respective single iteration.
6 . The system according to claim 5 , wherein the at least one automated processor is further configured to control the adaptation based on a change in the lower bound of policy performance as a selection criterion for a subsequent diverse set within a respective single iteration.
7 . The method according to claim 5 , wherein the at least one automated processor is further configured to control the adaptation is based on feedback of a system state of the artificial agent received after deploying a prior behavior policy within a respective single iteration.
8 . The system according to claim 1 , wherein the at least one automated processor is further configured to select the diverse set within a respective single iteration as the plurality of behavior policies generated based on prior feedback, having maximum differences from each other according to a Kullback-Leibler (KL) divergence measure.
9 . The system according to claim 1 , wherein the at least one automated processor is further configured to select the plurality of behavior policies of the diverse set within a respective iteration according to an aggregate group statistic.
10 . The system according to claim 1 , further comprising the artificial agent controlled according to a behavior policy;
wherein each iteration of policy improvement:
selects the diverse set of behavior policies from a space of stochastic behavior policies;
ensures that each respective behavior policy is safe and has a statistically expected return no worse than a lower threshold of policy performance;
maximizes behavior diversity of the diverse set according to a diversity metric;
updates a selection criterion; and
the artificial agent configured to control the system within the environment in accordance with a respective behavioral policy from the iteratively improved diverse set.
11 . The system according to claim 1 , wherein the at least one automated processor is further configured to employ importance sampling within a confidence interval to select the plurality of behavior policies within the diverse set for each iteration.
12 . The system according to claim 1 , wherein the at least one automated processor is further configured to collect the data set representing an environment in a first number of dimensions in each iteration, wherein the set of behavior policies have a second number of dimensions less than the first number of dimensions.
13 . The system according to claim 1 , wherein the at least one automated processor is further configured to update a statistically expected return, no worse than a lower bound of policy performance, between iterations.
14 . The system according to claim 1 , wherein the at least one automated processor is further configured, in each iteration, to obtain feedback from a system controlled by the artificial agent in accordance with the respective behavior policy, and to use the feedback to improve a computational model of the system which is predictive of future behavior of the system over a range of environmental conditions.
15 . The system according to claim 14 , wherein the at least one automated processor is further configured to implement a computational model of the system which is predictive of future behavior of the system over a multidimensional range of environmental conditions, based on a plurality of observations under different environmental conditions having a distribution, and to bias the diverse exploration strategy to select respective behavior policies within the set of behavior policies which selectively explore portions of the multidimensional range of environmental conditions.
16 . The system according to claim 1 , wherein the at least one automated processor is further configured to select the diverse set based on a predicted state of a system controlled by the artificial agent according to the respective behavior policy during deployment of the respective behavior policy.
17 . The system according to claim 1 , wherein the at least one automated processor is further configured to select the diverse set for assessment within each iteration to generate a maximum predicted statistical improvement in policy performance.
18 . A method of iterative policy improvement in reinforcement learning, comprising:
in each policy improvement iteration i, deploying a most recently confirmed set of policies to collect n trajectories uniformly distributed over the respective policies π i within the set of policies π i ∈ ; for each set of trajectories i collected from a respective policy π i , partition i and append to a training set of trajectories train and a testing set of trajectories test ; from train , generating a set of candidate policies and evaluating them using test ; confirming a subset of policies as meeting predetermined criteria; and if no new policies π i are confirmed, redeploying the current set of policies .
19 . The method according to claim 18 , further comprising, for each iteration:
defining a lower policy performance bound ρ.; and performing a t-test on normalized returns of test without importance sampling, treating the set of deployed policies as a mixture policy that generated test , wherein the set of policies are a set of conjugate policies generated as a byproduct of conjugate gradient descent.
20 . An apparatus for learning and deploying a set of behavior policies for an artificial agent, selected from a set of behavior policies, to control a system within an environment, comprising:
an input configured to receive data from operation of the system according to a respective behavior policy; at least one automated processor configured to iteratively generate sets of behavior policies based on prior received data which iteratively improve a behavior policy for each iteration of policy improvement, each behavior policy having a statistically expected return no worse than a lower bound of policy performance which excludes a portion of the set of behavior policies and which is ensured as being safe, employing a diverse exploration strategy which strives for behavior diversity in a space of stochastic policies by deploying a diverse set comprising a plurality of behavior policies during each iteration of policy improvement; and at least one output configured to control the system in accordance with a respective behavior policy of the set of behavior policies.Join the waitlist — get patent alerts
Track US2023169342A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.