Estimating depth for image and relative camera poses between images
Abstract
A computer implemented method of estimating depth for an image and relative camera poses between images in a video sequence, includes backwards warping the source image to generate a first reconstructed target image, and calculating an initial image reconstruction loss based on the target image and the first reconstructed target image. Forward warping the source depth map is performed to generate a second reconstructed target depth map, and an occlusion mask is generated based on the second reconstructed target depth map. The method further includes regularising the initial image reconstruction loss based on the generated occlusion mask. Thus, an occlusion aware method of image reconstruction is provided via a combination of forward and backward warping which identifies and masks occluded areas, and regularizes the image reconstruction loss.
Claims
exact text as granted — not AI-modified1 . A computer implemented method of estimating depth for an image and relative camera poses between images in a video sequence, comprising:
estimating a target depth map or a target image in a time series of two or more images; estimating a pose transformation from the target image to a source image, adjacent to the target image in the time series; backwards warping the source image to generate a first reconstructed target image, based on the pose transformation and the target depth map; calculating an initial image reconstruction loss, based on the target image and the first reconstructed target image; estimating a source depth map for the source image; forward warping the source depth map to generate a second reconstructed target depth map, based on the pose transformation and the source depth map; generating an occlusion mask by the second reconstructed target depth map, the occlusion map indicating one or more occluded areas of the target image; and regularising the initial image reconstruction loss based on the generated occlusion mask to generate a regularised image reconstruction loss.
2 . The method of claim 1 , wherein the estimating the target depth map and the source depth map uses a first neural network.
3 . The method of claim 1 , further comprising training the first neural network based on the regularised image reconstruction loss.
4 . The method of claim 1 , wherein the estimating the pose transformation uses a second neural network.
5 . The method of claim 4 , further comprising training the second neural network based on the regularised image reconstruction loss.
6 . The method of claim 1 , the backwards warping comprising:
projecting a plurality of target pixel locations of the target image into a 3D space, based on the target depth map and a set of camera intrinsic parameters; transforming positions of the projected pixel locations to the source image, based on the pose transformation; mapping pixel values of the source image onto corresponding target pixel locations; and generating the first reconstructed target image based on the mapped pixel values.
7 . The method of claim 6 , wherein mapping the pixel values of the source image onto the corresponding target pixel locations includes determining a pixel value using bilinear sampling of pixel values from adjacent pixel locations of the source image if a transformed target pixel location does not fall into an integer pixel location in the source image.
8 . The method of claim 1 , wherein the forward warping comprises:
projecting a plurality of depth values from the source image into a 3D space based on the source depth map and a set of camera intrinsic parameters; generating a pose transformation from the source image to the target image by reversing the pose transformation from the target image to the source image; transforming positions of the projected depth values, based on the pose transformation from the source image to the target image; and mapping the transformed depth values onto the second reconstructed target depth map based on the set of camera intrinsic parameters.
9 . The method of claim 8 , wherein the mapping the transformed depth values onto the second reconstructed target depth map includes determining a minimum depth value from an occluded set of depth values and discarding other depth values in the occluded set if the occluded set of depth values are mapped onto a single pixel location of the second reconstructed target depth map.
10 . A non-transitory computer-readable media storing computer instructions that configure at least one processor, upon execution of the instructions, to perform the following steps:
estimating a target depth map or a target image in a time series of two or more images; estimating a pose transformation from the target image to a source image, adjacent to the target image in the time series; backwards warping the source image to generate a first reconstructed target image, based on the pose transformation and the target depth map; calculating an initial image reconstruction loss, based on the target image and the first reconstructed target image; estimating a source depth map for the source image; forward warping the source depth map to generate a second reconstructed target depth map, based on the pose transformation and the source depth map; generating an occlusion mask by the second reconstructed target depth map, the occlusion map indicating one or more occluded areas of the target image; and regularising the initial image reconstruction loss based on the generated occlusion mask to generate a regularised image reconstruction loss.
11 . The computer-readable media of claim 10 , wherein the estimating the target depth map and the source depth map uses a first neural network.
12 . The computer-readable media of claim 10 , wherein the computer instructions further configure the at least one processor, upon execution of the instructions, to perform training the first neural network based on the regularised image reconstruction loss.
13 . The computer-readable media of claim 10 , wherein the estimating the pose transformation uses a second neural network.
14 . The computer-readable media of claim xx, wherein the computer instructions further configure the at least one processor, upon execution of the instructions, to perform training the second neural network based on the regularised image reconstruction loss.
15 . The computer-readable media of claim 10 , the backwards warping comprising:
projecting a plurality of target pixel locations of the target image into a 3D space, based on the target depth map and a set of camera intrinsic parameters; transforming positions of the projected pixel locations to the source image, based on the pose transformation; mapping pixel values of the source image onto corresponding target pixel locations; and generating the first reconstructed target image based on the mapped pixel values.
16 . The computer-readable media of claim 15 , wherein mapping the pixel values of the source image onto the corresponding target pixel locations includes determining a pixel value using bilinear sampling of pixel values from adjacent pixel locations of the source image if a transformed target pixel location does not fall into an integer pixel location in the source image.
17 . The computer-readable media of claim 10 , wherein the forward warping comprises:
projecting a plurality of depth values from the source image into a 3D space based on the source depth map and a set of camera intrinsic parameters; generating a pose transformation from the source image to the target image by reversing the pose transformation from the target image to the source image; transforming positions of the projected depth values, based on the pose transformation from the source image to the target image; and mapping the transformed depth values onto the second reconstructed target depth map based on the set of camera intrinsic parameters.
18 . The computer-readable media of claim 17 , wherein the mapping the transformed depth values onto the second reconstructed target depth map includes determining a minimum depth value from an occluded set of depth values and discarding other depth values in the occluded set if the occluded set of depth values are mapped onto a single pixel location of the second reconstructed target depth map.Join the waitlist — get patent alerts
Track US2023351624A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.