US2023376106A1PendingUtilityA1

Depth information based pose determination for mobile platforms, and associated systems and methods

Assignee: SZ DJI TECHNOLOGY CO LTDPriority: Dec 13, 2017Filed: Jul 31, 2023Published: Nov 23, 2023
Est. expiryDec 13, 2037(~11.4 yrs left)· nominal 20-yr term from priority
G06F 3/011G06T 7/593H04N 13/271G06V 10/426G06V 20/64G06V 40/103G06T 7/521H04N 2013/0081G06F 3/017G06F 3/0304
70
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method includes determining a depth range where a subject is likely to appear in a current depth map based on one or more previous depth maps of the environment, filtering the current depth map based on the depth range, to generate a reference depth map, identifying a plurality of candidate regions from the reference depth map, selecting a subset of the plurality of candidate regions, determining a main region from the subset of the plurality of candidate regions, associating the main region and one or more target regions, identifying the first pose component of the subject from a collective region, identifying the second pose component of the subject from the collective region, determining one or more vectors representing a spatial relationship between the identified first pose component and the identified second pose component, and controlling a movement of a movable object based on the one or more vectors.

Claims

exact text as granted — not AI-modified
I/We claim: 
     
         1 . A method comprising:
 determining a depth range where a subject is likely to appear in a current depth map of an environment based, at least in part, on one or more previous depth maps of the environment;   filtering the current depth map based, at least in part, on the depth range, to generate a reference depth map;   identifying a plurality of candidate regions from the reference depth map, a depth change of each of the plurality of candidate regions not exceeding a threshold, and the plurality of candidate regions being disconnected with each other;   selecting a subset of the plurality of candidate regions, a size of each candidate region in the subset of the plurality of candidate regions being within a threshold range, and the threshold range being determined based on an estimated size of a first pose component of the subject;   determining a main region from the subset of the plurality of candidate regions based, at least in part, on a position or a size corresponding to a second pose component of the subject;   generating a collective region by associating the main region with one or more target regions based, at least in part, on a relative position between the main region and the one or more target regions of the subset of the plurality of candidate regions, the one or more target regions being likely to correspond to one or more parts of the subject;   identifying the first pose component of the subject from the collective region;   identifying the second pose component of the subject from the collective region after the identified first pose component is identified from the collective region;   determining one or more vectors representing a spatial relationship between the identified first pose component and the identified second pose component; and   controlling a movement of a movable object based, at least in part, on the one or more vectors.   
     
     
         2 . The method of  claim 1 , further comprising:
 generating depth data based, at least in part, on obtained images captured by a stereo camera carried by the movable object.   
     
     
         3 . The method of  claim 2 , wherein:
 the depth data includes the current depth map calculated based on a disparity map or intrinsic parameters of the stereo camera.   
     
     
         4 . The method of  claim 2 , wherein the depth data includes at least one of unknown, invalid, or inaccurate depth information. 
     
     
         5 . The method of  claim 1 , wherein
 the movable object includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, a robot, a smart wearable device, a virtual reality (VR) head-mounted display, or an augmented reality (AR) head-mounted display.   
     
     
         6 . The method of  claim 1 , wherein the subject includes a human. 
     
     
         7 . The method of  claim 1 , wherein:
 the first pose component includes one of the one or more body parts of the subject; and   the second pose component includes another one of the one or more body parts of the subject.   
     
     
         8 . The method of  claim 1 , wherein:
 the first pose component includes a hand of the subject; and   the second pose component includes a torso of the subject.   
     
     
         9 . The method of  claim 1 , wherein:
 identifying the first pose component and the second pose component of the subject from the collective region includes detecting a portion of the collective region based, at least in part, on a measurement of depth.   
     
     
         10 . The method of  claim 9 , wherein:
 detecting the portion of the collective region includes detecting a portion of the collective region that is closest in depth to the movable object.   
     
     
         11 . A movable object comprising:
 a controller programmed to control the movable object, wherein the controller includes one or more processors configured to:
 determine a depth range where a subject is likely to appear in a current depth map of an environment based, at least in part, on one or more previous depth maps of the environment; 
 filter the current depth map based, at least in part, on the depth range, to generate a reference depth map; 
 identify a plurality of candidate regions from the reference depth map, a depth change of each of the plurality of candidate regions not exceeding a threshold, and the plurality of candidate regions being disconnected with each other; 
 select a subset of the plurality of candidate regions, a size of each candidate region in the subset of the plurality of candidate regions being within a threshold range, and the threshold range being determined based on an estimated size of a first pose component of the subject; 
 determine a main region from the subset of the plurality of candidate regions based, at least in part, on a position or a size corresponding to a second pose component of the subject; 
 generate a collective region by associating the main region with one or more target regions based, at least in part, on a relative position between the main region and the one or more target regions of the subset of the plurality of candidate regions, the one or more target regions being likely to correspond to one or more parts of the subject; 
 identify the first pose component of the subject from the collective region; 
 identify the second pose component of the subject from the collective region after the identified first pose component is identified from the collective region; 
 determine one or more vectors representing a spatial relationship between the identified first pose component and the identified second pose component; and 
 control a movement of the movable object based, at least in part, on the one or more vectors. 
   
     
     
         12 . The movable object of  claim 11 , further comprising:
 a stereo camera;   wherein the one or more processors are further configured to generate depth data based, at least in part, on obtained images captured by the stereo camera.   
     
     
         13 . The movable object of  claim 12 , wherein:
 the depth data includes the current depth map calculated based on a disparity map or intrinsic parameters of the stereo camera.   
     
     
         14 . The movable object of  claim 12 , wherein the depth data includes at least one of unknown, invalid, or inaccurate depth information. 
     
     
         15 . The movable object of  claim 11 , wherein
 the movable object includes at least one of an unmanned aerial vehicle (UAV), a manned aircraft, an autonomous car, a self-balancing vehicle, a robot, a smart wearable device, a virtual reality (VR) head-mounted display, or an augmented reality (AR) head-mounted display.   
     
     
         16 . The movable object of  claim 11 , wherein:
 the first pose component includes one of the one or more body parts of the subject; and   the second pose component includes another one of the one or more body parts of the subject.   
     
     
         17 . The movable object of  claim 11 , wherein:
 the first pose component includes a hand of the subject; and   the second pose component includes a torso of the subject.   
     
     
         18 . The movable object of  claim 11 , wherein:
 identifying the first pose component and the second pose component of the subject from the collective region includes detecting a portion of the collective region based, at least in part, on a measurement of depth.   
     
     
         19 . The movable object of  claim 18 , wherein:
 detecting the portion of the collective region includes detecting a portion of the collective region that is closest in depth to the movable object.   
     
     
         20 . A non-transitory computer-readable medium storing computer-executable instructions that, when executed, cause one or more processors associated with a movable object to perform actions, the actions comprising:
 determining a depth range where a subject is likely to appear in a current depth map of an environment based, at least in part, on one or more previous depth maps of the environment;   filtering the current depth map based, at least in part, on the depth range, to generate a reference depth map;   identifying a plurality of candidate regions from the reference depth map, a depth change of each of the plurality of candidate regions not exceeding a threshold, and the plurality of candidate regions being disconnected with each other;   selecting a subset of the plurality of candidate regions, a size of each candidate region in the subset of the plurality of candidate regions being within a threshold range, and the threshold range being determined based on an estimated size of a first pose component of the subject;   determining a main region from the subset of the plurality of candidate regions based, at least in part, on a position or a size corresponding to a second pose component of the subject;   generating a collective region by associating the main region with one or more target regions based, at least in part, on a relative position between the main region and the one or more target regions of the subset of the plurality of candidate regions, the one or more target regions being likely to correspond to one or more parts of the subject;   identifying the first pose component of the subject from the collective region;   identifying the second pose component of the subject from the collective region after the identified first pose component is identified from the collective region;   determining one or more vectors representing a spatial relationship between the identified first pose component and the identified second pose component; and   controlling a movement of the movable object based, at least in part, on the one or more vectors.

Join the waitlist — get patent alerts

Track US2023376106A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.