US2021227236A1PendingUtilityA1

Scalability of multi-directional video streaming

Assignee: APPLE INCPriority: Sep 14, 2018Filed: Apr 2, 2021Published: Jul 22, 2021
Est. expirySep 14, 2038(~12.1 yrs left)· nominal 20-yr term from priority
H04N 19/167H04N 21/44004H04N 21/21805H04N 19/597G09G 5/14H04N 19/176H04N 19/17H04N 19/29H04N 19/162H04N 19/103
56
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Aspects of the present disclosure provide techniques for reducing latency and improving image quality of a viewport extracted from multi-directional video communications. According to such techniques, first streams of coded video data are received from a source. The first streams include coded data for each of a plurality of tiles representing a multi-directional video, where each tile corresponding to a predetermined spatial region of the multi-directional video, and at least one tile of the plurality of tiles in the first streams contains a current viewport location at a receiver. The techniques include decoding the first streams and displaying the tile containing the current viewport location. When the viewport location at the receiver changes to include a new tile of the plurality of tiles, retrieving and decoding first streams for the new tile, displaying the decoded content for the changed viewport location, and transmitting the changed viewport location to the source.

Claims

exact text as granted — not AI-modified
1 .- 21 . (canceled) 
     
     
         22 . A video reception method, comprising:
 receiving a coded bitstream of multi-directional video of a scene including a first version of a frame in a first projection format and a second version of the frame in a second projection format;   first decoding the first version to produce a first decoded image in the first projection format;   second decoding the second version to produce a second decoded image in the second projection format;   converting the first decoded image from the first projection format to the second projection format;   combining the first decoded image in the second projection format with the second decoded image in the second projection format to produce a combined image in the second projection format; and   outputting the combined image as a decoded version of the frame.   
     
     
         23 . The method of  claim 22 , wherein:
 the first projection format is a equirectangular projection;   the second projection format is a cube map projection;   
     
     
         24 . The method of  claim 22 , wherein:
 the combined image represents a region of interest that correspond to a subset of the first projection format and a subset of the second projection format; and   pixels in the combined image are based on a weighted combination of corresponding pixels in the first decoded image with corresponding pixels in the second decoded image.   
     
     
         25 . The method of  claim 22 , wherein:
 the coded bitstream was encoded with a layered coding technique;   the first version is a base layer of the layered coding technique;   the second version is an interlayer prediction residual for an enhancement layer of the layered coding technique;   the converting predicts an enhancement layer output from the first version; and   the combining combines the predicted enhancement layer with the interlayer prediction residual to produce a decoded enhancement layer output.   
     
     
         26 . The method of  claim 25 , wherein:
 the base layer of the first version spatially includes the entire multi-directional scene; and   the enhancement layer of the second version includes a spatial region of interest that is a subset of the entire multi-directional scene.   
     
     
         27 . The method of  claim 25 , wherein:
 the first projection format is equirectangular projection and the base layer of the first version spatially includes the entire multi-directional scene; and   the second projection format is a cube map projection and the enhancement layer of the second version includes one face of the cube map projection and is a subset of the entire multi-directional scene.   
     
     
         28 . A video reception system, comprising:
 a receiver for receiving, from a source, a coded bitstream of multi-directional video of a scene including a first version of a frame in a first projection format and a second version of the frame in a second projection format;   a decoder for decoding the coded bitstream;   a controller to controlling the decoder to cause:
 first decoding the first version to produce a first decoded image in the first projection format; 
 second decoding the second version to produce a second decoded image in the second projection format; 
 converting the first decoded image from the first projection format to the second projection format; 
 combining the first decoded image in the second projection format with the second decoded image in the second projection format to produce a combined image in the second projection format; and 
 outputting the combined image as a decoded version of the frame. 
   
     
     
         29 . The method of  claim 28 , wherein:
 the first projection format is a equirectangular projection;   the second projection format is a cube map projection;   
     
     
         30 . The system of  claim 28 , wherein:
 the combined image represents a region of interest that correspond to a subset of the first projection format and a subset of the second projection format; and   pixels in the combined image are based on a weighted combination of corresponding pixels in the first decoded image with corresponding pixels in the second decoded image.   
     
     
         31 . The system of  claim 28 , wherein:
 the coded bitstream was encoded with a layered coding technique;   the first version is a base layer of the layered coding technique;   the second version is an interlayer prediction residual for an enhancement layer of the layered coding technique;   the converting predicts an enhancement layer output from the first version; and   the combining combines the predicted enhancement layer with the interlayer prediction residual to produce a decoded enhancement layer output.   
     
     
         32 . The system of  claim 31 , wherein:
 the base layer of the first version spatially includes the entire multi-directional scene; and   the enhancement layer of the second version includes a spatial region of interest that is a subset of the entire multi-directional scene.   
     
     
         33 . The system of  claim 31 , wherein:
 the first projection format is equirectangular projection and the base layer of the first version spatially includes the entire multi-directional scene; and   the second projection format is a cube map projection and the enhancement layer of the second version includes one face of the cube map projection and is a subset of the entire multi-directional scene.   
     
     
         34 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause:
 receiving a coded bitstream of multi-directional video of a scene including a first version of a frame in a first projection format and a second version of the frame in a second projection format;   first decoding the first version to produce a first decoded image in the first projection format;   second decoding the second version to produce a second decoded image in the second projection format;   converting the first decoded image from the first projection format to the second projection format;   combining the first decoded image in the second projection format with the second decoded image in the second projection format to produce a combined image in the second projection format; and   outputting the combined image as a decoded version of the frame.   
     
     
         35 . The computer readable medium of  claim 34 , wherein:
 the first projection format is a equirectangular projection;   the second projection format is a cube map projection;   
     
     
         36 . The computer readable medium of  claim 34 , wherein:
 the combined image represents a region of interest that correspond to a subset of the first projection format and a subset of the second projection format; and   pixels in the combined image are based on a weighted combination of corresponding pixels in the first decoded image with corresponding pixels in the second decoded image.   
     
     
         37 . The computer readable medium of  claim 34 , wherein:
 the coded bitstream was encoded with a layered coding technique;   the first version is a base layer of the layered coding technique;   the second version is an interlayer prediction residual for an enhancement layer of the layered coding technique;   the converting predicts an enhancement layer output from the first version; and   the combining combines the predicted enhancement layer with the interlayer prediction residual to produce a decoded enhancement layer output.   
     
     
         38 . The computer readable medium of  claim 37 , wherein:
 the base layer of the first version spatially includes the entire multi-directional scene; and   the enhancement layer of the second version includes a spatial region of interest that is a subset of the entire multi-directional scene.   
     
     
         39 . The computer readable medium of  claim 37 , wherein:
 the first projection format is equirectangular projection and the base layer of the first version spatially includes the entire multi-directional scene; and   the second projection format is a cube map projection and the enhancement layer of the second version includes one face of the cube map projection and is a subset of the entire multi-directional scene.

Join the waitlist — get patent alerts

Track US2021227236A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.