Image fusion for autonomous vehicle operation
Abstract
Devices, systems and methods for fusing scenes from real-time image feeds from on-vehicle cameras in autonomous vehicles to reduce redundancy of the information processed to enable real-time autonomous operation are described. One example of a method for improving perception in an autonomous vehicle includes receiving a plurality of cropped images, wherein each of the plurality of cropped images comprises one or more bounding boxes that correspond to one or more objects in a corresponding cropped image; identifying, based on the metadata in the plurality of cropped images, a first bounding box in a first cropped image and a second bounding box in a second cropped image, wherein the first and second bounding boxes correspond to a common object; and fusing the metadata corresponding to the common object from the first cropped image and the second cropped image to generate an output result for the common object.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method implemented by a processor disposed in a vehicle, the method comprising:
receiving images from one or more sensors installed in the vehicle; applying, on the images, a detection algorithm stored in the processor to identify a first bounding box in a first image and a second bounding box in a second image, wherein the first bounding box and the second bounding box are associated with metadata providing information corresponding to a common object; and generating an output including a fusion of the first image and the second image, the output having a size smaller than a sum of a size of the first image and a size of the second image.
2 . The method of claim 1 , further comprising:
fusing the first image and the second image based on the metadata.
3 . The method of claim 1 , further comprising, before the applying of the detection algorithm:
selecting and cropping one or more regions of interest in an image.
4 . The method of claim 1 , wherein the first image and the second image are from a same sensor.
5 . The method of claim 1 , wherein the first image and the second image are from different sensors facing a substantially similar direction.
6 . The method of claim 1 , wherein the common object is a vehicle, and wherein the metadata of a bounding box comprises at least one of a vehicle feature vector, a taillight signal detection result or a vehicle segmentation mask corresponding to the vehicle detected in the first bounding box or the second bounding box.
7 . The method of claim 1 , wherein the common object is a common vehicle, wherein the first image comprises a left taillight and a right taillight of the common vehicle, wherein the second image comprises exactly one taillight of the common vehicle, and wherein the output comprises a taillight signal detection that is a majority vote based on the left taillight, the right taillight and the exactly one taillight.
8 . The method of claim 1 , wherein the common object is a common vehicle, wherein the first image comprises a first vehicle segmentation mask corresponding to the common vehicle, wherein the second image comprises a second vehicle segmentation mask corresponding to the common vehicle, and wherein the output comprises a vehicle segmentation mask based on a convex combination of the first vehicle segmentation mask and the second vehicle segmentation mask.
9 . The method of claim 1 , wherein the applying the detection algorithm provides detection outputs having a unified focal plane.
10 . An apparatus implemented in a vehicle, comprising:
a plurality of sensors; a processor; and a memory with instructions thereon, wherein images are generated from one or more images captured by at least one of the plurality of sensors, wherein the instructions upon execution by the processor cause the processor to:
identify, based on metadata in the images, a first bounding box in a first image and a second bounding box in a second image, wherein the first bounding box and the second bounding box correspond to a common object;
generate an output including a fusion of the first image and the second image,
wherein a size of the output is smaller than a sum of a size of the first image and a size of the second image.
11 . The apparatus of claim 10 , wherein the metadata includes at least one of 2D or 3D detection results, a vehicle-type classification, a vehicle identification, a taillight signal detection results, or a vehicle segmentation mask.
12 . The apparatus of claim 10 , wherein the first image and the second image are from a same sensor.
13 . The apparatus of claim 10 , wherein the first image and the second image are from different sensors facing a substantially similar direction.
14 . The apparatus of claim 10 , wherein the common object is a vehicle, a pedestrian, a structure adjacent to roadways.
15 . A computer-readable storage medium having code stored thereon, the code, upon execution by one or more processors, causing the one or more processors to implement a method comprising:
receiving, from one or more sensors, images; applying, on the images, a detection algorithm stored in the one or more processors to identify a first bounding box in a first image and a second bounding box in a second image, wherein the first bounding box and the second bounding box are associated with metadata providing information corresponding to a common object; and generating an output including a fusion of the first image and the second image, the output having a size smaller than a sum of a size of the first image and a size of the second image.
16 . The computer-readable storage medium of claim 15 , wherein the metadata of the first bounding box comprises at least one of an object feature vector, a taillight signal detection result or an object segmentation mask corresponding to the common object detected in the first bounding box.
17 . The computer-readable storage medium of claim 16 , wherein the metadata of the first bounding box in the first image further comprises at least one of a camera pose, a focal length, a shutter speed or a field-of-view associated with a sensor that has generated the first image.
18 . The computer-readable storage medium of claim 16 , wherein the object feature vector comprises a color of an object or a make of the object.
19 . The computer-readable storage medium of claim 16 , wherein the object segmentation mask comprises one or more contours of an object.
20 . The computer-readable storage medium of claim 15 , wherein the images are generated by exactly one of the one or more sensors or two or more of the one or more sensors facing towards a substantially similar direction.Join the waitlist — get patent alerts
Track US2024046654A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.