US2023202031A1PendingUtilityA1

Machine learning control of object handovers

Assignee: NVIDIA CORPPriority: Jul 28, 2020Filed: Mar 1, 2023Published: Jun 29, 2023
Est. expiryJul 28, 2040(~14 yrs left)· nominal 20-yr term from priority
G06V 40/107B25J 9/16B25J 9/1612B25J 9/1697G06V 20/64G06V 20/30G06T 2207/10028G06T 7/50G06N 3/08G06N 20/00G06N 5/041G06N 3/045B25J 9/163G06T 7/73G06T 2207/20081G06T 2207/20084G05B 2219/40202G05B 2219/40609G05B 2219/40559G05B 2219/40613G05B 2219/39543G05B 2219/39487G05B 2219/39271G05B 2219/39546G06V 40/11G06V 10/82G06V 10/764
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A robotic control system directs a robot to take an object from a human grasp by obtaining an image of a human hand holding an object, estimating the pose of the human hand and the object, and determining a grasp pose for the robot that will not interfere with the human hand. In at least one example, a depth camera is used to obtain a point cloud of the human hand holding the object. The point cloud is provided to a deep network that is trained to generate a grasp pose for a robotic gripper that can take the object from the human's hand without pinching or touching the human's fingers.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A processor, comprising one or more circuits to:
 generate a plurality of grasp poses that allow a robot to grasp an object; and   select, based at least in part on a pose of a hand, a target grasp pose from the plurality of grasp poses that does not interfere with the hand.   
     
     
         2 . The processor of  claim 1 , wherein the one or more circuits are to:
 obtain an image comprising the hand gripping the object;   segment the image to identify a portion of the image that represents the pose of the hand and a second portion of the image that represents a pose of the object; and   generate the plurality of grasp poses based, at least in part, on the segmented image.   
     
     
         3 . The processor of  claim 1 , wherein the one or more circuits are to:
 use one or more images to perform classification of the hand grasping the object and to determine a position of the hand; and   generate the plurality of grasp poses based, at least in part, on the classification and the position of the hand.   
     
     
         4 . The processor of  claim 1 , wherein the one or more circuits are to:
 use one or more images to identify the pose of the hand gripping the object;   generate the plurality of grasp poses based, at least in part, on the pose of the hand gripping the object;   use the selected target grasp pose from the plurality of grasp poses to adjust a position of the robot to grasp the object; and   cause the robot to use the target grasp pose that does not interfere with the hand to grasp the object.   
     
     
         5 . The processor of  claim 1 , wherein the one or more circuits are to:
 generate a motion plan that comprises information to cause the robot to avoid contact between the robot and the hand; and   use the motion plan and the target grasp pose to grasp the object.   
     
     
         6 . The processor of  claim 1 , wherein the one or more circuits are to:
 use the selected target grasp pose to grasp the object; and   generate a motion plan to cause the robot to drop the object, wherein the motion plan includes information to cause the robot to move to a position to drop the object that avoids a collision.   
     
     
         7 . The processor of  claim 1 , wherein the one or more circuits are to use one or more neural networks to generate the plurality of grasp poses, wherein the one or more neural networks are trained by one or more images comprising one or more hands gripping the object. 
     
     
         8 . A system, comprising one or more processors to:
 generate a plurality of grasp poses that allow a robot to grasp an object; and   select, based at least in part on a pose of a hand, a target grasp pose from the plurality of grasp poses that does not interfere with the hand.   
     
     
         9 . The system of  claim 8 , wherein the one or more processors are to:
 use one or more images to generate a human grasp data set; and   train one or more neural networks using the generated human grasp data set to estimate the pose of the hand.   
     
     
         10 . The system of  claim 8 , wherein the one or more processors are to:
 use one or more neural networks to classify the pose of the hand grasping the object; and   generate a plan for the robot to grasp the object from the hand based, at least in part, on the classification.   
     
     
         11 . The system of  claim 8 , wherein the interference comprises the robot being in contact with the hand. 
     
     
         12 . The system of  claim 8 , wherein the one or more processors are to adjust the target grasp pose based, at least in part, on a motion of the hand. 
     
     
         13 . The system of  claim 8 , wherein the one or more processors are to use the target grasp pose to cause the robot to grasp the object from a hand of another robot. 
     
     
         14 . The system of  claim 8 , wherein the one or more processors are to move the robot to a position to avoid colliding with the hand if a distance between the robot and hand exceeds a threshold. 
     
     
         15 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
 generate a plurality of grasp poses that allow a robot to grasp an object; and   select, based at least in part on a pose of a hand, a target grasp pose from the plurality of grasp poses that does not interfere with the hand.   
     
     
         16 . The non-transitory machine-readable medium of  claim 15 , wherein the instructions, if performed by one or more processors, cause the one or more processors to adjust the motion of the robot to avoid colliding with the hand after grasping the object from the hand. 
     
     
         17 . The non-transitory machine-readable medium of  claim 15 , wherein the object is being held in an open palm of the hand or held by two or more fingers of the hand. 
     
     
         18 . The non-transitory machine-readable medium of  claim 15 , wherein the instructions, if performed by one or more processors, cause the one or more processors to generate the plurality of grasp poses by one or more neural networks, wherein the one or more neural networks are trained by one or more segmented images comprising the hand gripping the object. 
     
     
         19 . The non-transitory machine-readable medium of  claim 15 , wherein the instructions, if performed by one or more processors, cause the one or more processors to:
 calculate one or more distances between the robot and the hand; and   use the calculated one or more distances to cause the robot to move to a position to avoid a collision with the hand.   
     
     
         20 . The non-transitory machine-readable medium of  claim 15 , wherein the instructions, if performed by one or more processors, cause the one or more processors to cause the robot to use the selected target grasp pose to grasp the object from an appendage, wherein the appendage comprises a body part of a human, an animal, or another robot.

Join the waitlist — get patent alerts

Track US2023202031A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.