Method and System for Encoding a 3D Scene
Abstract
A computer-implemented method for encoding a scene volume includes: (a) identifying features of a scene volume that are within a camera perspective range with respect to a default camera perspective; (b) converting the identified features into rendered features; and (c) sorting the rendered features into a plurality of scene layers, each including corresponding depth, color, and transparency maps for the respective rendered features. Further, (a), (b), and (c) may be repeated, operating on temporally ordered scene volumes, to produce and output a sequence encoding a video. Corresponding systems and non-transitory computer-readable media are disclosed for encoding a 3D scene and for decoding an encoded 3D scene. Efficient compression, transmission, and playback of video describing a 3D scene can be enabled, including for virtual reality displays with updates based on a changing perspective of a user viewer for variable-perspective playback.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for encoding a scene volume, the method comprising:
(a) identifying features of a scene volume that are within a camera perspective range with respect to a default camera perspective; (b) converting the identified features into rendered features; and (c) sorting the rendered features into a plurality of scene layers, wherein each respective scene layer of the plurality of scene layers includes, for respective rendered features that are sorted into the respective scene layer: (i) a corresponding depth map for the respective rendered features, the depth map based on depth information from the scene volume, (ii) a corresponding color map for the respective rendered features, and (iii) a corresponding transparency map for the respective rendered features.
2 . The method of claim 1 , wherein sorting the rendered features into the plurality of scene layers includes sorting by depth of the rendered features.
3 . The method of claim 1 , further including writing the plurality of scene layers to a non-transitory computer-readable medium.
4 . The method of claim 1 , further including constructing the scene volume using sequential color images captured by a camera.
5 . The method of claim 1 , further including constructing the scene volume using a single frame from a camera.
6 . The method of claim 5 , wherein constructing the scene volume includes inferring, using a machine learning model, one or more features that are occluded from the single frame.
7 . The method of claim 1 , further including repeating (a), (b), and (c), for a respective plurality of temporally ordered scene volumes, to produce a sequence of temporally ordered pluralities of scene layers.
8 . The method of claim 7 , further including encoding the sequence of pluralities of scene layers to create a compressed sequence.
9 . The method of claim 7 , further including packaging each plurality of scene layers into a series of image files, and then compressing the series of image files into a respective video file corresponding to the respective plurality of scene layers.
10 . A system for encoding a scene volume, the system comprising:
one or more processors configured to:
(a) identify features of a scene volume that are within a camera perspective range with respect to a default camera perspective;
(b) convert the identified features into rendered features; and
(c) sort the rendered features into a plurality of scene layers, wherein each respective scene layer of the plurality of scene layers includes, for respective rendered features that are sorted into the respective scene layer: (i) a corresponding depth map for the respective rendered features, the depth map based on depth information from the scene volume, (ii) a corresponding color map for the respective rendered features, and (iii) a corresponding transparency map for the respective rendered features.
11 . The system of claim 10 , further including a non-transitory computer-readable medium, and wherein the one or more processors are further configured to output the plurality of scene layers for storage in the non-transitory computer-readable medium.
12 . The system of claim 11 , wherein the one or more processors are further configured to: receive a plurality of temporally ordered scene volumes; to identify, convert, and sort according to (a), (b), and (c), respectively; and to output a sequence of pluralities of scene layers into the non-transitory computer medium.
13 . A computer-implemented method for generating a scene volume, the method comprising:
(a) assigning depth to respective features of a scene volume based on respective depth maps of respective scene layers, the respective scene layers being of a plurality of scene layers of an encoded scene volume; (b) assigning color to the respective features of the scene volume based on respective color maps of the respective scene layers; and (c) assigning transparency to the respective features of the scene volume based on respective transparency maps of the respective scene layers.
14 . The method of claim 13 , further comprising receiving the plurality of scene layers.
15 . The method of claim 13 , further comprising creating a rendered perspective from the scene volume.
16 . The method of claim 13 , further including generating a plurality of temporally ordered scene volumes by repeating (a), (b), and (c) for a plurality of respective, temporally ordered, pluralities of scene layers.
17 . A system for generating a scene volume, the system comprising:
one or more processors configured to: (a) assign depth to respective features of a scene volume based on respective depth maps of respective scene layers, the respective scene layers being of a plurality of scene layers of an encoded scene volume; (b) assign color to the respective features of the scene volume based on respective color maps of the respective scene layers; and (c) assign transparency to the respective features of the scene volume based on respective transparency maps of the scene layers.
18 . One or more non-transitory, computer-readable media comprising instructions that, when executed by one or more processors, cause a device to:
(a) identify features of a scene volume that are within a camera perspective range with respect to a default camera perspective; (b) convert the identified features into rendered features; and (c) sort the rendered features into a plurality of scene layers, wherein each respective scene layer of the plurality of scene layers includes, for respective rendered features that are sorted into the respective scene layer: (i) a corresponding depth map for the respective rendered features, the depth map based on depth information from the scene volume, (ii) a corresponding color map for the respective rendered features, and (iii) a corresponding transparency map for the respective rendered features.
19 . One or more non-transitory, computer-readable media comprising instructions that, when executed by one or more processors, cause a device to:
(a) assign depth to respective features of a scene volume based on respective depth maps of respective scene layers, the respective scene layers being of a plurality of scene layers of an encoded scene volume; (b) assign color to the respective features of the scene volume based on respective color maps of the respective scene layers; and (c) assign transparency to the respective features of the scene volume based on respective transparency maps of the scene layers.Join the waitlist — get patent alerts
Track US2022353530A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.