Gesture recognition using depth images
Abstract
Methods, apparatuses, and articles associated with gesture recognition using depth images are disclosed herein. In various embodiments, an apparatus may include a face detection engine configured to determine whether a face is present in one or more gray images of respective image frames generated by a depth camera, and a hand tracking engine configured to track a hand in one or more depth images generated by the depth camera. The apparatus may further include a feature extraction and gesture inference engine configured to extract features based on results of the tracking by the hand tracking engine, and infer a hand gesture based at least in part on the extracted features. Other embodiments may also be disclosed and claimed.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . (canceled)
3 . (canceled)
4 . (canceled)
5 . (canceled)
6 . (canceled)
7 . (canceled)
8 . (canceled)
9 . (canceled)
10 . (canceled)
11 . (canceled)
12 . (canceled)
13 . A method comprising:
tracking, by a computing apparatus, a hand in selected respective regions of one or more depth images generated by a depth camera, wherein the selected respective regions are size-wise smaller than the respective one or more depth images; and inferring a hand gesture, by the computing device, based at least in part on a result of the tracking; wherein tracking comprises determining location measures of the hand for the depth images.
14 . The method of claim 13 , wherein determining location measures of the hand for the depth images comprises determining a pair of (x, y) coordinates for a center of the hand, using mean-shift filtering that uses gradients of probabilistic density.
15 . The method of claim 13 , wherein inferring comprises extracting one or more features for the selected respective regions, based at least in part on a result of the tracking, and inferring a hand gesture based at least in part on the extracted one or more features.
16 . The method of claim 15 , wherein extracting one or more features comprises extracting one or more of an eccentricity measure, a compactness measure, an orientation measure, a rectangularity measure, a horizontal center measure, a vertical center measure, a minimum bounding box angle measure, a minimum bounding box width-to-height ratio measure, a difference between left-and-right measure, or a difference between up-and-down measure.
17 . The method of claim 13 , wherein inferring a gesture comprises inferring one of an open gesture, a fist gesture, a thumb up gesture, a thumb down gesture, a thumb left gesture or a thumb right gesture.
18 . (canceled)
19 . A method comprising:
extracting, by a computing apparatus, one or more features from respective regions of depth images of image frames generated by a depth camera; and inferring a gesture, by the computing apparatus, based at least in part on the one or more features extracted from the depth images; wherein extracting one or more features comprises extracting one or more of an eccentricity measure, a compactness measure, an orientation measure, a rectangularity measure, a horizontal center measure, a vertical center measure, a minimum bounding box angle measure, a minimum bounding box width-to-height ratio measure, a difference between left-and-right measure, or a difference between up-and-down measure.
20 . The method of claim 19 , wherein extracting one or more features from respective regions of depth images comprises extracting one or more features from respective regions of depth images denoted as containing a hand.
21 . The method of claim 19 wherein inferring a gesture comprises inferring one of an open gesture, a fist gesture, a thumb up gesture, a thumb down gesture, a thumb left gesture or a thumb right gesture.
22 . (canceled)
23 . A computer-readable non-transitory storage medium, comprising
a plurality of programming instructions stored in the storage medium, and configured to cause an apparatus, in response to execution of the programming instructions by the apparatus, to perform operations including: tracking a hand in selected respective regions of one or more depth images generated by a depth camera, wherein the selected respective regions are size-wise smaller than the respective one or more depth images; and inferring a hand gesture, based at least in part on a result of the tracking; wherein tracking comprises determining location measures of the hand for the depth images.
24 . The storage medium of claim 23 , wherein determining location measures of the hand for the depth images comprises determining a pair of (x, y) coordinates for a center of the hand, using mean-shift filtering that uses gradients of probabilistic density.
25 . The storage medium of claim 23 , wherein inferring comprises extracting one or more features for the selected respective regions, based at least in part on a result of the tracking, and inferring a hand gesture based at least in part on the extracted one or more features.
26 . The storage medium of claim 25 , wherein extracting one or more features comprises extracting one or more of an eccentricity measure, a compactness measure, an orientation measure, a rectangularity measure, a horizontal center measure, a vertical center measure, a minimum bounding box angle measure, a minimum bounding box width-to-height ratio measure, a difference between left-and-right measure, or a difference between up-and-down measure.
27 . The storage medium of claim 23 , wherein inferring a gesture comprises inferring one of an open gesture, a fist gesture, a thumb up gesture, a thumb down gesture, a thumb left gesture or a thumb right gesture.
28 . The storage medium of claim 23 , wherein the operations further comprise determining whether a face is present in the one or more depth images' corresponding one or more gray images of respective image frames generated by a depth camera.
29 . The storage medium of claim 28 , wherein determining whether a face is present comprises analyzing the one or more gray images using a Haar-Cascade model.
30 . The storage medium of claim 28 , wherein the operations further comprise determining a measure of a distance between the face and the camera, using the one or more depth images.
31 . An apparatus, comprising:
a tracking engine to track a hand in selected respective regions of one or more depth images generated by a depth camera, wherein the selected respective regions are size-wise smaller than the respective one or more depth images; and an inference engine coupled with the tracking engine to infer a hand gesture based at least in part on a result of the tracking; wherein to track a hand comprises to determine location measures of the hand for the depth images.
32 . The apparatus of claim 31 , wherein to determine location measures of the hand for the depth images comprises to determine a pair of (x, y) coordinates for a center of the hand, using mean-shift filtering that uses gradients of probabilistic density.
33 . The apparatus of claim 31 , wherein to infer comprises to extract one or more features for the selected respective regions, based at least in part on a result of the tracking, and to infer a hand gesture based at least in part on the extracted one or more features.
34 . The apparatus of claim 33 , wherein to extract one or more features comprises to extract one or more of an eccentricity measure, a compactness measure, an orientation measure, a rectangularity measure, a horizontal center measure, a vertical center measure, a minimum bounding box angle measure, a minimum bounding box width-to-height ratio measure, a difference between left-and-right measure, or a difference between up-and-down measure.
35 . The apparatus of claim 31 , wherein to infer a gesture comprises to infer one of an open gesture, a fist gesture, a thumb up gesture, a thumb down gesture, a thumb left gesture or a thumb right gesture.
36 . An apparatus comprising:
an extraction engine to extract one or more features from respective regions of depth images of image frames generated by a depth camera; and an inference engine coupled with the extraction engine to infer a gesture, based at least in part on the one or more features extracted from the depth images; wherein to extract one or more features comprises to extract one or more of an eccentricity measure, a compactness measure, an orientation measure, a rectangularity measure, a horizontal center measure, a vertical center measure, a minimum bounding box angle measure, a minimum bounding box width-to-height ratio measure, a difference between left-and-right measure, or a difference between up-and-down measure.
37 . The apparatus of claim 36 , wherein to extract one or more features from respective regions of depth images comprises to extract one or more features from respective regions of depth images denoted as containing a hand.
38 . The apparatus of claim 36 wherein to infer a gesture comprises to infer one of an open gesture, a fist gesture, a thumb up gesture, a thumb down gesture, a thumb left gesture or a thumb right gesture.Join the waitlist — get patent alerts
Track US2014300539A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.