Machine learning device, robot system, and machine learning method for learning object picking operation
Abstract
A machine learning device that learns an operation of a robot for picking up, by a hand unit, any of a plurality of workpieces placed in a random fashion, including a bulk-loaded state, includes a state variable observation unit that observes a state variable representing a state of the robot, including data output from a three-dimensional measuring device that obtains a three-dimensional map for each workpiece, an operation result obtaining unit that obtains a result of a picking operation of the robot for picking up the workpiece by the hand unit, and a learning unit that learns a manipulated variable including command data for commanding the robot to perform the picking operation of the workpiece, in association with the state variable of the robot and the result of the picking operation, upon receiving output from the state variable observation unit and output from the operation result obtaining unit.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An estimation method, comprising:
obtaining, by at least one processor, data in relation to a plurality of objects including at least one of
information in relation to shapes of the plurality of objects, or
information after processing the information in relation to the shapes of the plurality of objects; and
causing, by the at least one processor, a neural network to generate information for picking up one of the plurality of objects by inputting the data in relation to the plurality of objects into the neural network.
2 . The estimation method according to claim 1 , wherein
the information after processing the information in relation to the shapes includes at least one of
position information of the plurality of objects,
orientation information of the plurality of objects, or
image information of the plurality of objects.
3 . The estimation method according to claim 1 , wherein
the information in relation to the shapes of the plurality of objects includes at least one of
image information of the plurality of objects,
three-dimensional position information of the plurality of objects, or
distance information from a measuring device to surfaces of the plurality of objects.
4 . The estimation method according to claim 1 , wherein
the neural network is updated by reinforcement learning using a reward calculated based on information in relation to a picking operation of an object.
5 . The estimation method according to claim 4 , wherein
the information in relation to the picking operation of the object includes at least one of
success or failure of the picking operation of the object,
a number of times of successes of picking operations of objects,
a time taken for picking up or transporting the object,
a force acting on a hand unit picking up or transporting the object,
an achievement level of a post-process after the picking operation of the object,
a change in state of the object, or
energy for picking up or transporting the object.
6 . The estimation method according to claim 4 , wherein
the information in relation to the picking operation of the object includes information for changing positions of a plurality of objects.
7 . The estimation method according to claim 4 , wherein
the neural network is a value function in the reinforcement learning.
8 . The estimation method according to claim 7 , wherein
the value function represents a value of control information of a robot picking up or transporting the object.
9 . The estimation method according to claim 1 , wherein
the neural network is updated to minimize an error calculated based on a label for picking up an object and an output of the neural network.
10 . The estimation method according to claim 9 , wherein
the neural network outputs at least one of position information of the object or information in relation to a success rate of picking up the object.
11 . The estimation method according to claim 1 , further comprising:
determining, by the at least one processor, whether the information for picking up the one of the plurality of objects is abnormal.
12 . The estimation method according to claim 1 , wherein
the neural network is updated by using data obtained from simulations.
13 . The estimation method according to claim 1 , wherein
the information for picking up the one of the plurality of objects includes at least one of
robot control information,
position information of a hand picking up or transporting the one of the plurality of objects,
orientation information of the hand,
take-out direction information of the hand,
position information of the one of the plurality of objects,
success rate information of picking up an object, or
control information of a measuring device.
14 . An estimation device, comprising:
at least one memory; and at least one processor coupled to the at least one memory and configured to:
obtain data in relation to a plurality of objects including at least one of
information in relation to shapes of the plurality of objects, or
information after processing the information in relation to the shapes of the plurality of objects; and
estimate information for picking up one of the plurality of objects by inputting the data in relation to the plurality of objects into a neural network.
15 . A learning method, comprising:
obtaining, by at least one processor, data in relation to a plurality of objects including at least one of
information in relation to shapes of the plurality of objects, or
information after processing the information in relation to the shapes of the plurality of objects; and
learning, by the at least one processor, a neural network to output information for picking up one of the plurality of objects by inputting the data in relation to the plurality of objects into the neural network.
16 . The learning method according to claim 15 , wherein
the information after processing the information in relation to the shapes includes at least one of
position information of the plurality of objects,
orientation information of the plurality of objects, or
image information of the plurality of objects.
17 . The learning method according to claim 15 , wherein
the information in relation to the shapes of the plurality of objects includes at least one of
image information of the plurality of objects,
three-dimensional position information of the plurality of objects, or
distance information from a measuring device to surfaces of the plurality of objects.
18 . The learning method according to claim 15 , further comprising:
updating the neural network by reinforcement learning using a reward calculated based on information in relation to a picking operation of an object.
19 . The learning method according to claim 18 , wherein
the information in relation to the picking operation of the object includes at least one of
success or failure of the picking operation of the object,
a number of times of successes of picking operations of objects,
a time taken for picking up or transporting the object,
a force acting on a hand unit picking up or transporting the object,
an achievement level of a post-process after the picking operation of the object,
a change in state of the object, or
energy for picking up or transporting the object.
20 . The learning method according to claim 18 , wherein
the information in relation to the picking operation of the object includes information for changing positions of a plurality of objects.
21 . The learning method according to claim 18 , wherein
the neural network is a value function in the reinforcement learning.
22 . The learning method according to claim 21 , wherein
the value function represents a value of control information of a robot picking up or transporting the object.
23 . The learning method according to claim 15 , further comprising:
updating the neural network to minimize an error calculated based on a label for picking up an object and an output of the neural network.
24 . The learning method according to claim 23 , wherein
the neural network outputs at least one of position information of the object or information in relation to a success rate of picking up the object.
25 . The learning method according to claim 15 , further comprising:
determining whether the information for picking up the one of the plurality of objects is abnormal.
26 . The learning method according to claim 15 , further comprising:
updating the neural network by using data obtained from simulations.
27 . The learning method according to claim 15 , wherein
the information for picking up the one of the plurality of objects includes at least one of
robot control information,
position information of a hand picking up or transporting the one of the plurality of objects,
orientation information of the hand,
take-out direction information of the hand,
position information of the one of the plurality of objects,
success rate information of picking up an object, or
control information of a measuring device.
28 . A learning device, comprising:
at least one memory; and at least one processor coupled to the at least one memory and configured to:
obtain data in relation to a plurality of objects including at least one of
information in relation to shapes of the plurality of objects, or
information after processing the information in relation to the shapes of the plurality of objects; and
learn a neural network to output information for picking up one of the plurality of objects from the neural network by inputting the data in relation to the plurality of objects into the neural network.Join the waitlist — get patent alerts
Track US2023321837A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.