Reinforcement Learning Method, Non-Transitory Computer Readable Recording Medium, Reinforcement Learning Device and Molding Machine
Abstract
A reinforcement learning method of a learning machine including a first agent adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device and a second agent having a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent, comprises: adjusting the manufacture condition searched by the first agent that is performing reinforcement learning, using the observation data and the functional model or the functional approximator of the second agent; calculating reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted; and performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated.
Claims
exact text as granted — not AI-modified1 . A reinforcement learning method for a learning machine including
a first agent adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device and a second agent having a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent, the reinforcement learning method comprising: adjusting the manufacture condition searched by the first agent that is performing reinforcement learning, using the observation data and the functional model or the functional approximator of the second agent; calculating reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted; and performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated.
2 . The reinforcement learning method according to claim 1 , comprising:
calculating a search range of the manufacture condition using the observation data and the functional model or the functional approximator of the second agent, and in a case where the manufacture condition searched by the first agent that is performing reinforcement learning falls out of the search range calculated, changing the manufacture condition searched to the manufacture condition falling within the search range.
3 . The reinforcement learning method according to claim 2 , comprising:
acquiring a threshold for calculating the search range of the manufacture condition using the observation data and the functional model or the functional approximator of the second agent, and calculating the search range of the manufacture condition using the threshold acquired, the observation data and the functional model or the functional approximator of the second agent.
4 . The reinforcement learning method according to claim 2 , comprising, in a case where the manufacture condition searched by the first agent that is performing reinforcement learning falls out of a predetermined search range, changing the manufacture condition searched to the manufacture condition falling within the predetermined search range and the search range calculated.
5 . The reinforcement learning method according to claim 1 , comprising, in a case where the manufacture condition searched by the first agent is adjusted by the second agent, calculating the reward data by adding a minus reward in accordance with a deviation degree of the first agent from a search range.
6 . The reinforcement learning method according to claim 1 , wherein the manufacturing device is a molding machine.
7 . The reinforcement learning method according to claim 6 , wherein
the manufacturing device is an injection molding machine, the manufacture condition includes an in-mold resin temperature, a nozzle temperature, a cylinder temperature, a hopper temperature, a mold clamping force, an injection speed, an injection acceleration, an injection peak pressure, an injection stroke, a cylinder-tip resin pressure, a reverse flow preventive ring seating state, a holding pressure switching pressure, a holding pressure switching speed, a holding pressure switching position, a holding pressure completion position, a cushion position, a metering back pressure, a metering torque, a metering completion position, a screw retreat speed, a cycle time, a mold closing time, an injection time, a pressure holding time, a metering time and a mold opening time, and the reward data is data calculated based on observation data of the injection molding machine or a defect degree of a molded product manufactured by the injection molding machine.
8 . A non-transitory computer readable recording medium storing a computer program causing a computer to perform reinforcement learning on a learning machine including
a first agent adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device and a second agent having a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent, the computer program causing the computer to execute processing of: adjusting the manufacture condition searched by the first agent that is performing reinforcement learning using the observation data and the functional model or the functional approximator of the second agent; calculating reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted; and performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated.
9 . A reinforcement learning device performing reinforcement learning on a learning machine adjusting a manufacture condition of a manufacturing device based on observation data obtained by observing a state of the manufacturing device, wherein
the learning machine comprising a first agent that adjusts the manufacture condition of the manufacturing device based on the observation data; a second agent that has a functional model or a functional approximator representing a relationship between the observation data and the manufacture condition in a different way from the first agent; an adjustment unit that adjusts the manufacture condition searched by the first agent that is performing reinforcement learning, using the observation data and the functional model or the functional approximator of the second agent; and a reward calculation unit that calculates reward data in accordance with a state of a product manufactured by the manufacturing device under the manufacture condition adjusted, the learning machine performing reinforcement learning on the first agent and the second agent based on the observation data and the reward data calculated by the reward calculation unit.
10 . A molding machine comprising:
the reinforcement learning device according to claim 9 , and a manufacturing device operated using the manufacture condition adjusted by the first agent.Join the waitlist — get patent alerts
Track US2024227266A9 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.