Cross-modal sensor data alignment
Abstract
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for determining an alignment between cross-modal sensor data. In one aspect, a method comprises: obtaining (i) an image that characterizes a visual appearance of an environment, and (ii) a point cloud comprising a collection of data points that characterizes a three-dimensional geometry of the environment; processing each of a plurality of regions of the image using a visual embedding neural network to generate a respective embedding of each of the image regions; processing each of a plurality of regions of the point cloud using a shape embedding neural network to generate a respective embedding of each of the point cloud regions; and identifying a plurality of region pairs using the embeddings of the image regions and the embeddings of the point cloud regions.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method performed by one or more data processing apparatus for aligning multi-modal sensor data, the method comprising:
obtaining multi-modal sensor data characterizing an environment, wherein the multi-modal sensor data comprises: (i) first sensor data generated by a first sensor modality, and (ii) second sensor data generated by a second sensor modality, wherein the second sensor modality is different than the first sensor modality; processing each of a plurality of regions of the first sensor data using a first embedding neural network that is specific to the first sensor modality to generate a respective region embedding of each of the plurality of regions of the first sensor data; processing each of a plurality of regions of the second sensor data using a second embedding neural network that is specific to the second sensor modality to generate a respective region embedding of each of the plurality of regions of the second sensor data; determining a plurality of similarity scores, wherein each similarity score measures a similarity between a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data; and identifying a plurality of region embedding pairs that collectively define an alignment of the first sensor data and the second sensor data based on the plurality of similarity scores, wherein each region embedding pair comprises a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data.
22 . The method of claim 21 , wherein the first sensor modality is an imaging modality and the first sensor data comprises an image that characterizes a visual appearance of the environment.
23 . The method of claim 21 , wherein the second sensor modality is a surveying sensor modality and the second sensor data comprises a point cloud, wherein the point cloud includes a collection of data points that characterize a three-dimensional geometry of the environment, wherein each data point defines a respective three-dimensional spatial position of a point on a surface in the environment.
24 . The method of claim 23 , wherein the second sensor data is captured by a lidar sensor or a radar sensor.
25 . The method of claim 24 , wherein the second sensor data is captured by a lidar sensor, and each data point in the point cloud additionally defines a strength of a reflection of a pulse of light that was transmitted by the lidar sensor and that reflected from the point on the surface of the environment at the three-dimensional spatial position defined by the data point.
26 . The method of claim 21 , wherein the first sensor data and the second sensor data are captured by sensors mounted on a vehicle.
27 . The method of claim 21 , further comprising:
using the alignment of the first sensor data and the second sensor data to determine whether a first sensor that captured the first sensor data and a second sensor that captured the second sensor data are accurately calibrated.
28 . The method of claim 21 , further comprising:
obtaining data defining a position of an object in the first sensor data; and identifying a corresponding position of the object in the second sensor data based on: (i) the position of the object in the first sensor data, and (ii) the alignment of the first sensor data and the second sensor data.
29 . The method of claim 21 , further comprising:
generating fused sensor data by fusing the first sensor data and the second sensor data using the alignment of the first sensor data and the second sensor data; and processing the fused sensor data using a neural network to generate a neural network output.
30 . The method of claim 29 , wherein the neural network output comprises data identifying positions of objects in the environment.
31 . The method of claim 21 , wherein the plurality of regions of the first sensor data cover the first sensor data.
32 . The method of claim 21 , wherein the plurality of regions of the second sensor data cover the second sensor data.
33 . The method of claim 21 , wherein the plurality of region embedding pairs are identified using a greedy nearest neighbor matching algorithm.
34 . The method of claim 21 , wherein the first embedding neural network and the second embedding neural network are jointly trained using a triplet loss objective function or a contrastive loss objective function.
35 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for aligning multi-modal sensor data, the operations comprising:
obtaining multi-modal sensor data characterizing an environment, wherein the multi-modal sensor data comprises: (i) first sensor data generated by a first sensor modality, and (ii) second sensor data generated by a second sensor modality, wherein the second sensor modality is different than the first sensor modality; processing each of a plurality of regions of the first sensor data using a first embedding neural network that is specific to the first sensor modality to generate a respective region embedding of each of the plurality of regions of the first sensor data; processing each of a plurality of regions of the second sensor data using a second embedding neural network that is specific to the second sensor modality to generate a respective region embedding of each of the plurality of regions of the second sensor data; determining a plurality of similarity scores, wherein each similarity score measures a similarity between a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data; and identifying a plurality of region embedding pairs that collectively define an alignment of the first sensor data and the second sensor data based on the plurality of similarity scores, wherein each region embedding pair comprises a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data.
36 . The non-transitory computer storage media of claim 35 , wherein the first sensor modality is an imaging modality and the first sensor data comprises an image that characterizes a visual appearance of the environment.
37 . The non-transitory computer storage media of claim 35 , wherein the second sensor modality is a surveying sensor modality and the second sensor data comprises a point cloud, wherein the point cloud includes a collection of data points that characterize a three-dimensional geometry of the environment, wherein each data point defines a respective three-dimensional spatial position of a point on a surface in the environment.
38 . A system comprising:
one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for aligning multi-modal sensor data, the operations comprising: obtaining multi-modal sensor data characterizing an environment, wherein the multi-modal sensor data comprises: (i) first sensor data generated by a first sensor modality, and (ii) second sensor data generated by a second sensor modality, wherein the second sensor modality is different than the first sensor modality; processing each of a plurality of regions of the first sensor data using a first embedding neural network that is specific to the first sensor modality to generate a respective region embedding of each of the plurality of regions of the first sensor data; processing each of a plurality of regions of the second sensor data using a second embedding neural network that is specific to the second sensor modality to generate a respective region embedding of each of the plurality of regions of the second sensor data; determining a plurality of similarity scores, wherein each similarity score measures a similarity between a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data; and identifying a plurality of region embedding pairs that collectively define an alignment of the first sensor data and the second sensor data based on the plurality of similarity scores, wherein each region embedding pair comprises a region embedding of a respective region of the first sensor data and a region embedding of a respective region of the second sensor data.
39 . The system of claim 38 , wherein the first sensor modality is an imaging modality and the first sensor data comprises an image that characterizes a visual appearance of the environment.
40 . The system of claim 38 , wherein the second sensor modality is a surveying sensor modality and the second sensor data comprises a point cloud, wherein the point cloud includes a collection of data points that characterize a three-dimensional geometry of the environment, wherein each data point defines a respective three-dimensional spatial position of a point on a surface in the environment.Join the waitlist — get patent alerts
Track US2022076082A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.