Automated optimization of a mass policy collectively performed for objects in two or more states and a direct policy performed in each state
Abstract
An information processing apparatus that optimizes a policy in a transition model in which the number of targeted objects in each state transits according to the policy includes a cost constraint acquisition unit configured to acquire a cost constraint that constrains a total cost of the policy; a mass policy setting unit configured to set the number of objects targeted by a mass policy in each state, based on the predefined number of objects to belong to each state and a reach rate at which the mass policy reaches to an object, with respect to the mass policy collectively executed for the object in two or more states; and a processing unit configured to assume the reach rate of the mass policy as a variable of an optimization and maximize an objective function based on a total reward in a whole period while satisfying the cost constraint.
Claims
exact text as granted — not AI-modified1 . An information processing method of optimizing a policy in a transition model in which the number of objects in each state transits according to the policy, the method being executed by a computer, the method comprising:
a cost constraint acquisition stage of acquiring a cost constraint that constrains a total cost of the policy; a mass policy setting stage of setting the number of objects targeted by a mass policy in each state, based on the predefined number of objects to belong to each state and a reach rate at which the mass policy reaches to an object, with respect to the mass policy collectively executed for the object in two or more states; and a processing stage of assuming the reach rate of the mass policy as a variable of an optimization and maximizing an objective function based on a total reward in a whole period while satisfying the cost constraint.
2 . The information processing method of claim 1 , wherein, in the mass policy setting stage, the number of objects targeted by the mass policy in each of the two or more states is set, based on the predefined number of objects to belong to each state and the reach rate common in the two or more states, with respect to the mass policy collectively executed for the object in the two or more states.
3 . The information processing method of claim 1 , wherein:
in the mass policy setting stage, the number of objects targeted by the mass policy in each state at each timing is set, based on the predefined number of objects in each state at each timing and the reach rate at which the mass policy reaches to the object, with respect to the mass policy; and in the processing stage, the reach rate in each timing with respect to the mass policy is assumed as a variable of an optimization, policy distribution in each state at each timing with respect to a direct policy executed every state is assumed as a variable of an optimization, and the objective function is maximized while the cost constraint is satisfied.
4 . The information processing method of claim 3 , wherein:
in the processing stage, policy distribution about the direct policy without the mass policy is assumed as a variable of an optimization and policy distribution that maximizes the objective function is calculated; in the mass policy setting stage, the predefined number of objects in the mass policy is set and the number of objects targeted by the mass policy in each state is set, based on a result acquired by maximizing the objective function excluding the mass policy; and in the processing stage, the reach rate in each timing with respect to the mass policy is assumed as the variable of the optimization, the policy distribution in each state at each timing with respect to the direct policy executed every state is assumed as the variable of the optimization, and the objective function is maximized while the cost constraint is satisfied.
5 . The information processing method of claim 1 , wherein:
in the mass policy setting stage, the predefined number of objects in the mass policy is set and the number of objects targeted by the mass policy in each state is set, based on a result acquired by maximizing the objective function while satisfying the cost constraint; and in the processing stage, the reach rate in each timing with respect to the mass policy is assumed as the variable of the optimization, the policy distribution in each state at each timing with respect to the direct policy executed every state is assumed as the variable of the optimization, and processing to maximize the objective function while satisfying the cost constraint is performed again.Join the waitlist — get patent alerts
Track US2015294350A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.