Tracking users across image frames using fingerprints obtained from image analysis
Abstract
Systems and methods are disclosed herein for tracking a vulnerable road user (VRU) regardless of occlusion. In an embodiment, the system captures a series of images including the VRU, and inputs each of the images into a detection model. The system receives a bounding box for each of the series of images of the VRU as output from the detection model. The system inputs each bounding box into a multi-task model, and receives as output from the multi-task model an embedding for each bounding box. The system determines, using the embeddings for each bounding box across the series of images, an indication of which of the embeddings correspond to the VRU.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
capturing a series of images comprising a plurality of humans, a human of the plurality of humans at least partially occluded in at least some images of the series of images; determining one or more respective bounding boxes for each respective image of the series of images, each respective bounding box for a respective human within the respective image; inputting each bounding box into a multi-task model; receiving as output from the multi-task model an embedding for each bounding box, the embedding produced from a shared layer of the multi-task model, the multi-task model comprising the shared layer and a plurality of branches each trained to predict a different activity, wherein the shared layer is trained using backpropagation from the plurality of branches; and determining, using the embeddings for each bounding box across the series of images, an indication of which of the embeddings correspond to the human as opposed to a different human of the plurality of humans despite partial occlusion of the human.
2 . The method of claim 1 , wherein capturing the series of images comprises receiving images captured by a camera installed on a vehicle, and wherein auxiliary data captured by sensors installed on the vehicle is received with the images.
3 . The method of claim 2 , further comprising:
receiving coordinates in the context of each of the series of images of each bounding box, wherein determining the indication of which of the embeddings correspond to the human comprises using the coordinates in addition to the embeddings.
4 . The method of claim 3 , wherein determining the indication of which of the embeddings correspond to the human further comprises using the auxiliary data.
5 . The method of claim 1 , wherein each respective embedding acts as a fingerprint that tracks its respective human without assigning an identity to the respective human.
6 . The method of claim 1 , wherein each human of the plurality of humans is a vulnerable road user (VRU).
7 . The method of claim 1 , wherein determining the indication of which of the embeddings correspond to the human further comprises receiving, as part of the output, a confidence score corresponding to a confidence that each given embedding corresponds to its indicated cluster.
8 . A non-transitory computer-readable medium comprising memory with instructions encoded thereon, the instructions causing one or more processors to perform operations when executed, the instructions comprising instructions to:
capture a series of images comprising a plurality of humans, a human of the plurality of humans at least partially occluded in at least some images of the series of images; determine one or more respective bounding boxes for each respective image of the series of images, each respective bounding box for a respective human within the respective image; input each bounding box into a multi-task model; receive as output from the multi-task model an embedding for each bounding box, the embedding produced from a shared layer of the multi-task model, the multi-task model comprising the shared layer and a plurality of branches each trained to predict a different activity, wherein the shared layer is trained using backpropagation from the plurality of branches; and determine, using the embeddings for each bounding box across the series of images, an indication of which of the embeddings correspond to the human as opposed to a different human of the plurality of humans despite partial occlusion of the human.
9 . The non-transitory computer-readable medium of claim 8 , wherein the instructions to capture the series of images comprises receiving images captured by a camera installed on a vehicle, and wherein auxiliary data captured by sensors installed on the vehicle is received with the images.
10 . The non-transitory computer-readable medium of claim 9 , the instructions further comprising instructions to:
receive coordinates in the context of each of the series of images of each bounding box, wherein determining the indication of which of the embeddings correspond to the human comprises using the coordinates in addition to the embeddings.
11 . The non-transitory computer-readable medium of claim 10 , wherein the instructions to determine the indication of which of the embeddings correspond to the human further comprise instructions to use the auxiliary data.
12 . The non-transitory computer-readable medium of claim 8 , wherein each respective embedding acts as a fingerprint that tracks its respective human without assigning an identity to the respective human.
13 . The non-transitory computer-readable medium of claim 8 , wherein each human of the plurality of humans is a vulnerable road user (VRU).
14 . The non-transitory computer-readable medium of claim 8 , wherein the instructions to determine the indication of which of the embeddings correspond to the human further comprise instructions to receive, as part of the output, a confidence score corresponding to a confidence that each given embedding corresponds to its indicated cluster.
15 . A system comprising:
a non-transitory computer-readable medium comprising memory with instructions encoded thereon; and one or more processors that, when executing the instructions, are caused to perform operations, the operations comprising:
capturing a series of images comprising a plurality of humans, a human of the plurality of humans at least partially occluded in at least some images of the series of images;
determining one or more respective bounding boxes for each respective image of the series of images, each respective bounding box for a respective human within the respective image;
inputting each bounding box into a multi-task model;
receiving as output from the multi-task model an embedding for each bounding box, the embedding produced from a shared layer of the multi-task model, the multi-task model comprising the shared layer and a plurality of branches each trained to predict a different activity, wherein the shared layer is trained using backpropagation from the plurality of branches; and
determining, using the embeddings for each bounding box across the series of images, an indication of which of the embeddings correspond to the human as opposed to a different human of the plurality of humans despite partial occlusion of the human.
16 . The system of claim 15 , wherein capturing the series of images comprises receiving images captured by a camera installed on a vehicle, and wherein auxiliary data captured by sensors installed on the vehicle is received with the images.
17 . The system of claim 16 , the operations further comprising:
receiving coordinates in the context of each of the series of images of each bounding box, wherein determining the indication of which of the embeddings correspond to the human comprises using the coordinates in addition to the embeddings.
18 . The system of claim 17 , wherein determining the indication of which of the embeddings correspond to the human further comprises using the auxiliary data.
19 . The system of claim 15 , wherein each respective embedding acts as a fingerprint that tracks its respective human without assigning an identity to the respective human.
20 . The system of claim 15 , wherein each human of the plurality of humans is a vulnerable road user (VRU).Join the waitlist — get patent alerts
Track US2023343062A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.