US2025310561A1PendingUtilityA1
Variable resolution variable frame rate video coding using neural networks
Est. expiryMay 11, 2042(~15.8 yrs left)· nominal 20-yr term from priority
G06F 3/013H04N 19/42H04N 19/33H04N 19/172H04N 19/136G06N 3/0455H04N 19/597H04N 19/31
49
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods for encoding video, and for decoding video at an arbitrary temporal and/or spatial resolution. The techniques use a scene representation neural network that, in implementations, is configured to represent frames of a 2D or 3D video as a 3D model encoded in the parameters of the neural network.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method of decoding video, comprising:
obtaining encoded video data for a video comprising a plurality of video frames, wherein the video frames comprise a sequence of sets of video frames, each set of video frames comprising the video frames between a respective pair of key frames, wherein the encoded video data for each set of video frames comprises parameters of a scene representation neural network that encodes the video frames between the respective pair of key frames, and wherein the scene representation neural network is configured to receive a representation of a frame time defining a time between the respective pair of key frames, and to process the representation of the frame time to generate a scene representation output for rendering an image that depicts a scene encoded by the parameters of the scene representation neural network at the frame time; and for each set of video frames in the encoded video data: processing a representation of each of a set of frame times between the respective pair of key frames for the set of video frames, using the scene representation neural network configured with the parameters for the set of video frames, to generate the scene representation output for each of the frame times; and rendering a set of image frames, one for each of the frame times, using the scene representation output for each of the frame times, wherein the set of image frames provides the decoded video.
2 . The method of claim 1 wherein the scene representation neural network is further configured to receive a representation of a viewing direction, and to process the representation of the frame time and the representation of the viewing direction to generate the scene representation output for rendering the image;
the method further comprising rendering each of the image frames by:
determining a viewing direction for each of a plurality of pixels in the image frame, wherein the viewing direction for a pixel corresponds to a direction of a ray into the scene from the pixel;
for each pixel in the image frame, processing the representation of the frame time and the representation of the viewing direction for the pixel, using the scene representation neural network to generate the scene representation output for the pixel and the frame time; and
rendering the image frame for the frame time using the scene representation outputs for the pixels of the image frame.
3 . The method of claim 2 , wherein the parameters of the scene representation neural network encode a representation of the scene over a three dimensional spatial volume, and wherein the scene representation neural network is further configured to receive a representation of a spatial location in the scene, and to process the representation of the frame time, the representation of the viewing direction, and the representation of the spatial location to generate the scene representation output, and wherein the scene representation output defines a light level emitted from the spatial location along the viewing direction and an opacity at the spatial location; and
wherein rendering each of the image frames comprises, for each pixel in the image frame: determining a plurality of spatial locations along the ray into the scene from the pixel; for each of the spatial locations, processing the representation of the frame time, the representation of the viewing direction, and the representation of the spatial location, using the scene representation neural network, to generate the scene representation output, wherein the scene representation output defines the light level emitted from the spatial location along the viewing direction and the opacity at the spatial location; and combining, for the spatial locations along the ray, the light level emitted from the spatial location along the viewing direction and the opacity at the spatial location, to determine a pixel value for the pixel in the image frame.
4 . The method of claim 2 , wherein the set of image frames comprise image frames defined on a concave 2D surface, and wherein the viewing direction for a pixel corresponds to a direction of a ray outwards from a point of view for the decoded video that is within the three dimensional spatial volume.
5 . The method of claim 2 , wherein rendering one of the image frames further comprises:
determining one or more increased spatial resolution areas of the image frame to be rendered at a higher spatial resolution than one or more other areas of the image frame; and rendering the one or more increased spatial resolution areas at the higher spatial resolution by rendering an increased area density of pixels for the one or more increased spatial resolution areas, rendering the increased area density of pixels comprising determining an increased number of viewing directions for the increased spatial resolution areas, the increased number of viewing directions corresponding to an increased density of rays into the scene compared to the one or more other areas of the image frame, the increased density of rays into the scene defining a reduced angular difference between the corresponding viewing directions.
6 . The method of claim 5 , wherein determining the one or more increased spatial resolution areas of the image frame comprises:
rendering the image frame at a first spatial resolution lower than the higher spatial resolution; determining a level of spatial detail in regions of the image frame rendered at the first spatial resolution; and determining the one or more increased spatial resolution areas as one or more areas that include a region of the image frame rendered at the first spatial resolution that has a higher level of spatial detail than another region of the image frame rendered at the first spatial resolution.
7 . The method of claim 6 , further comprising:
determining a level of spatial detail in regions of the one or more increased spatial resolution areas of the image frame; determining one or more further increased spatial resolution areas as one or more areas that include a region of the image frame rendered at the higher spatial resolution that has a higher level of spatial detail than another region of the image frame rendered at the higher spatial resolution; and rendering the one or more further increased spatial resolution areas at a further increased spatial resolution that is higher than the higher spatial resolution.
8 . The method of claim 5 wherein determining the one or more increased spatial resolution areas of the image frame comprises:
obtaining a gaze direction for an observer of the image frame;
determining the increased spatial resolution area of the image frame, using the gaze direction, as a part of the image frame to which the observer is directing their gaze.
9 . The method of claim 1 , further comprising:
determining a frame rate for the video; and determining the number of video frames in each set of video frames dependent upon the frame rate.
10 . The method of claim 1 , comprising determining the frame times dependent upon a metric of a rate of change of content of the video, such that a time interval between successive image frames of the decoded video is decreased when the metric indicates an increased rate of change.
11 . The method of claim 10 , wherein determining the frame times comprises:
rendering a first set of image frames, one for each of a first set of frame times; in response to the metric, determining one or more additional frame times temporally between successive frame times of the first set of frame times; and rendering one or more additional image frames corresponding to the additional frame times by processing representations of the additional frame times using the scene representation neural network.
12 . The method of claim 11 , wherein rendering one or more additional image frames corresponding to the additional frame times comprises rendering only part of the additional image frames that is determined by the metric to have an increased rate of change.
13 . The method of claim 1 , comprising:
obtaining a gaze direction for an observer of the image frames; and rendering part of one or more additional image frames corresponding to additional frame times temporally between successive frame times of the set of frame times, by processing representations of the additional frame times using the scene representation neural network; wherein the part of the one or more additional image frames comprises at least a part to which the observer is directing their gaze.
14 . The method of claim 5 , wherein determining the one or more increased spatial resolution areas of the image frame further comprises:
determining partial image frame times for one or more additional partial image frames, wherein each partial image frame comprises the one or more increased spatial resolution areas, and wherein the partial image frame times define an increased temporal resolution for the increased spatial resolution areas; and rendering the additional partial image frames by processing the partial image frame times using the using the scene representation neural network, wherein the processing comprises: processing representations of the partial image frame times and the representations of the viewing directions for pixels of the additional partial image frames using the scene representation neural network to generate the scene representation output, and rendering, for each partial image frame time, the respective additional partial image frame using the scene representation outputs for the pixels of the additional partial image frame.
15 . (canceled)
16 . A computer-implemented method of encoding video, comprising:
obtaining source video, the source video comprising a sequence of sets of source video frames, each set of source video frames comprising the source video frames between a respective pair of source video key frames; encoding the source video to obtain encoded video data, wherein encoding the source video comprises, for each set of source video frames between a respective pair of source video key frames: training an encoder scene representation neural network using each of the source video frames of the source video between the respective pair of source video key frames, to generate an encoder scene representation output for rendering an image that depicts a scene at a respective source frame time of the source video frame, wherein the scene is encoded by parameters of the encoder scene representation neural network; and storing or transmitting the encoded video data.
17 . The method of claim 16 wherein the encoder scene representation neural network is configured to receive a representation of the source frame time, and to process the representation of the source frame time to generate the encoder scene representation output for rendering an image that depicts a scene encoded by the parameters of the encoder scene representation neural network at the source frame time; and wherein the training comprises, for each of the source video frames:
processing the representation of the source frame time using the encoder scene representation neural network to generate the encoder scene representation output for the source frame time;
rendering the image that depicts the scene at the source frame time using the encoder scene representation output for the source frame time; and
updating the parameters of the encoder scene representation neural network using an objective function that characterizes an error between the source video frame and the rendered the image that depicts the scene at the source frame time.
18 . (canceled)
19 . A system comprising:
one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for decoding a video, the operations comprising: obtaining encoded video data for a video comprising a plurality of video frames, wherein the video frames comprise a sequence of sets of video frames, each set of video frames comprising the video frames between a respective pair of key frames, wherein the encoded video data for each set of video frames comprises parameters of a scene representation neural network that encodes the video frames between the respective pair of key frames, and wherein the scene representation neural network is configured to receive a representation of a frame time defining a time between the respective pair of key frames, and to process the representation of the frame time to generate a scene representation output for rendering an image that depicts a scene encoded by the parameters of the scene representation neural network at the frame time; and for each set of video frames in the encoded video data: processing a representation of each of a set of frame times between the respective pair of key frames for the set of video frames, using the scene representation neural network configured with the parameters for the set of video frames, to generate the scene representation output for each of the frame times; and rendering a set of image frames, one for each of the frame times, using the scene representation output for each of the frame times, wherein the set of image frames provides the decoded video.
20 . (canceled)
21 . The system of claim 19 wherein the scene representation neural network is further configured to receive a representation of a viewing direction, and to process the representation of the frame time and the representation of the viewing direction to generate the scene representation output for rendering the image;
the method further comprising rendering each of the image frames by:
determining a viewing direction for each of a plurality of pixels in the image frame, wherein the viewing direction for a pixel corresponds to a direction of a ray into the scene from the pixel;
for each pixel in the image frame, processing the representation of the frame time and the representation of the viewing direction for the pixel, using the scene representation neural network to generate the scene representation output for the pixel and the frame time; and
rendering the image frame for the frame time using the scene representation outputs for the pixels of the image frame.
22 . The system of claim 21 , wherein the parameters of the scene representation neural network encode a representation of the scene over a three dimensional spatial volume, and wherein the scene representation neural network is further configured to receive a representation of a spatial location in the scene, and to process the representation of the frame time, the representation of the viewing direction, and the representation of the spatial location to generate the scene representation output, and wherein the scene representation output defines a light level emitted from the spatial location along the viewing direction and an opacity at the spatial location; and
wherein rendering each of the image frames comprises, for each pixel in the image frame: determining a plurality of spatial locations along the ray into the scene from the pixel; for each of the spatial locations, processing the representation of the frame time, the representation of the viewing direction, and the representation of the spatial location, using the scene representation neural network, to generate the scene representation output, wherein the scene representation output defines the light level emitted from the spatial location along the viewing direction and the opacity at the spatial location; and combining, for the spatial locations along the ray, the light level emitted from the spatial location along the viewing direction and the opacity at the spatial location, to determine a pixel value for the pixel in the image frame.
23 . The system of claim 21 , wherein the set of image frames comprise image frames defined on a concave 2D surface, and wherein the viewing direction for a pixel corresponds to a direction of a ray outwards from a point of view for the decoded video that is within the three dimensional spatial volume.Join the waitlist — get patent alerts
Track US2025310561A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.