Neural Super-sampling for Real-time Rendering
Abstract
In one embodiment, a method includes receiving a pair of stereo images having a resolution lower than a target resolution, generating an initial first feature map for a first image of the pair based on first channels associated with the first image and generating an initial second feature map for a second image of the pair based on second channels associated with the second image, generating a first feature map based on combining the first channels with the initial first feature map, generating a second feature map based on combining the second channels with the initial second feature map, up-sampling the first feature map and the second feature map to the target resolution, warping the up-sampled second feature map, and generating a reconstructed image corresponding to the first image having the target resolution based on the up-sampled first feature map and the up-sampled and warped second feature map.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising, by one or more computing systems:
receiving a pair of stereo images having a resolution lower than a target resolution; generating (a) an initial first feature map for a first image of the pair of stereo images based on one or more first channels associated with the first image and (b) an initial second feature map for a second image of the pair of stereo images based on one or more second channels associated with the second image; generating a first feature map based on combining the one or more first channels with the initial first feature map; generating a second feature map based on combining the one or more second channels with the initial second feature map; up-sampling the first feature map and the second feature map to the target resolution; warping the up-sampled second feature map; and generating a reconstructed image corresponding to the first image having the target resolution based on the up-sampled first feature map and the up-sampled and warped second feature map.
2 . The method of claim 1 , wherein the first or second image comprises an RGB image with depth information.
3 . The method of claim 2 , further comprising:
converting the RGB image with depth information to a YCbCr image.
4 . The method of claim 1 , wherein generating the first or second feature map is based on one or more convolutional neural networks.
5 . The method of claim 1 , wherein each of the initial first feature map and the initial second feature map is based on a first number of channels, and wherein each of the first feature map and the second feature map is based on a second number of channels.
6 . The method of claim 1 , wherein up-sampling the first feature map and the second feature map to the target resolution is based on zero up-sampling, wherein the zero up-sampling comprises:
assigning each input pixel of each of the first feature map and the second feature map to its corresponding pixel at the target resolution; and setting all missing pixels around the input pixel as zeros.
7 . The method of claim 1 , wherein warping the up-sampled second feature map is based on a motion estimation associated with the pair of stereo images.
8 . The method of claim 7 , wherein the pair of stereo images are received from a client device, wherein the method further comprises determining the motion estimation based on a head motion detected by the client device, comprising:
identifying a motion vector based on the head motion; and resizing the motion vector to the target resolution based on bilinear up-sampling.
9 . The method of claim 7 , wherein warping the up-sampled second feature map comprises using the motion estimation with bilinear interpolation during warping.
10 . The method of claim 1 , wherein up-sampling the first feature map to the target resolution is based on information associated with the second image.
11 . The method of claim 1 , further comprising:
inputting the up-sampled first feature map and the up-sampled and warped second feature map to a feature reweighting module, wherein the feature reweighting module is based on one or more convolutional neural networks.
12 . The method of claim 11 , further comprising:
generating, by the feature weighting module, a pixel-wise weighting map for the up-sampled and warped second feature map; and multiplying the pixel-wise weighting map with the up-sampled and warped second feature map to generate a reweighted feature map for the second image.
13 . The method of claim 12 , wherein generating the reconstructed image corresponding to the first image comprises:
combining the up-sampled first feature map and the reweighted feature map for the second frame.
14 . The method of claim 1 , wherein generating the reconstructed image corresponding to the first image is based on a machine-learning model, wherein the machine-learning model is based on a convolutional neural network with one or more skip connections.
15 . The method of claim 1 , wherein the first image is captured by a first camera, wherein the second image is captured by a second camera.
16 . The method of claim 15 , wherein warping the up-sampled second feature map comprises:
warping the up-sampled second feature map of the second image to a viewpoint of the first camera.
17 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
receive a pair of stereo images having a resolution lower than a target resolution; generate (a) an initial first feature map for a first image of the pair of stereo images based on one or more first channels associated with the first image and (b) an initial second feature map for a second image of the pair of stereo images based on one or more second channels associated with the second image; generate a first feature map based on combining the one or more first channels with the initial first feature map; generate a second feature map based on combining the one or more second channels with the initial second feature map; up-sample the first feature map and the second feature map to the target resolution; warp the up-sampled second feature map; and generate a reconstructed image corresponding to the first image having the target resolution based on the up-sampled first feature map and the up-sampled and warped second feature map.
18 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
receive a pair of stereo images having a resolution lower than a target resolution; generate (a) an initial first feature map for a first image of the pair of stereo images based on one or more first channels associated with the first image and (b) an initial second feature map for a second image of the pair of stereo images based on one or more second channels associated with the second image; generate a first feature map based on combining the one or more first channels with the initial first feature map; generate a second feature map based on combining the one or more second channels with the initial second feature map; up-sample the first feature map and the second feature map to the target resolution; warp the up-sampled second feature map; and generate a reconstructed image corresponding to the first image having the target resolution based on the up-sampled first feature map and the up-sampled and warped second feature map.Join the waitlist — get patent alerts
Track US2022277421A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.