Prediction precision improvements video coding
Abstract
Devices, systems and methods for digital video coding, which includes inter prediction with refinement, are described. An exemplary method of video processing includes determining to use, for a conversion between a current block of a video and a bitstream representation of the video, a first linear optimization model for the conversion using a first coding mode, the first linear optimization model being derived from a second linear optimization model that is used for the conversion using a second coding mode, and performing, based on the determining, the conversion. Another exemplary method of video processing includes determining to use, for a conversion between a current block of a video and a bitstream representation of the video, a gradient value computation algorithm for a bi-directional optical flow tool, and performing, based on the determining, the conversion.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of processing video data, comprising:
determining, for a conversion between a current block of a video and a bitstream of the video, two corresponding regions for a sub-block of the current block, wherein the two corresponding regions are in two reference pictures of the current block respectively; deriving a sum of absolute difference (SAD) based on partial samples of the two corresponding regions; determining a first prediction mode is applied to the sub-block in response to the SAD being larger than or equal to a first threshold, wherein the first prediction mode is an optical flow-based inter prediction mode; and deriving gradient values in different directions based on prediction samples of the sub-block and an arithmetic shifting operation with a shift value S, and wherein S is an integer and S is not equal to 4.
2 . The method of claim 1 , wherein S is equal to 6.
3 . The method of claim 1 , wherein samples in one row of every N rows in each of the two corresponding regions are used to derive the SAD, and N is an integer larger than 1.
4 . The method of claim 1 , wherein positions of the partial samples of the two corresponding regions are predetermined.
5 . The method of claim 1 , wherein the method further comprising:
deriving a horizontal motion offset and a vertical motion offset based on the gradient values; deriving a prediction refinement based on the horizontal motion offset, the vertical motion offsets, and the gradient values; deriving final prediction samples based on a sum of the prediction refinement and the prediction samples, wherein a clipping operation is applied to the sum of the prediction refinement and the prediction samples; and performing the conversion based on the final prediction samples.
6 . The method of claim 5 , wherein the clipping operation is within a range [minPred, maxPred], and wherein the maxPred is based on a sample bit-depth of the current block, the minPred and the maxPred are integers.
7 . The method of claim 1 , wherein the prediction samples are derived based on a filtering operation and a padding process.
8 . The method of claim 7 , wherein only an 8-tap interpolation filter is used in the filtering operation.
9 . The method of claim 1 , wherein the conversion includes encoding the current block into the bitstream.
10 . The method of claim 1 , wherein the conversion includes decoding the current block from the bitstream.
11 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:
determine, for a conversion between a current block of a video and a bitstream of the video, two corresponding regions for a sub-block of the current block, wherein the two corresponding regions are in two reference pictures of the current block respectively; derive a sum of absolute difference (SAD) based on partial samples of the two corresponding regions; determine a first prediction mode is applied to the sub-block in response to the SAD being larger than or equal to a first threshold, wherein the first prediction mode is an optical flow-based inter prediction mode; and derive gradient values in different directions based on prediction samples of the sub-block and an arithmetic shifting operation with a shift value S, and wherein S is an integer and S is not equal to 4.
12 . The apparatus of claim 11 , wherein S is equal to 6.
13 . The apparatus of claim 11 , wherein samples in one row of every N rows in each of the two corresponding regions are used to derive the SAD, and N is an integer larger than 1.
14 . The apparatus of claim 11 , wherein positions of the partial samples of the two corresponding regions are predetermined.
15 . The apparatus of claim 11 , wherein the instructions upon execution by the processor, further cause the processor to:
derive a horizontal motion offset and a vertical motion offset based on the gradient values; derive a prediction refinement based on the horizontal motion offset, the vertical motion offsets, and the gradient values; derive final prediction samples based on a sum of the prediction refinement and the prediction samples, wherein a clipping operation is applied to the sum of the prediction refinement and the prediction samples; and perform the conversion based on the final prediction samples.
16 . The apparatus of claim 15 , wherein the clipping operation is within a range [minPred, maxPred], and wherein the maxPred is based on a sample bit-depth of the current block, the minPred and the maxPred are integers.
17 . The apparatus of claim 11 , wherein the prediction samples are derived based on a filtering operation and a padding process.
18 . The apparatus of claim 17 , wherein only an 8-tap interpolation filter is used in the filtering operation.
19 . A non-transitory computer-readable recording medium storing a bitstream of a video which is generated by a method performed by a video processing apparatus, wherein the method comprises:
determining, for a conversion between a current block of a video and a bitstream of the video, two corresponding regions for a sub-block of the current block, wherein the two corresponding regions are in two reference pictures of the current block respectively; deriving a sum of absolute difference (SAD) based on partial samples of the two corresponding regions; determining a first prediction mode is applied to the sub-block in response to the SAD being larger than or equal to a first threshold, wherein the first prediction mode is an optical flow-based inter prediction mode; and deriving gradient values in different directions based on prediction samples of the sub-block and an arithmetic shifting operation with a shift value S, and wherein S is an integer and S is not equal to 4.
20 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:
determine, for a conversion between a current block of a video and a bitstream of the video, two corresponding regions for a sub-block of the current block, wherein the two corresponding regions are in two reference pictures of the current block respectively; derive a sum of absolute difference (SAD) based on partial samples of the two corresponding regions; determine a first prediction mode is applied to the sub-block in response to the SAD being larger than or equal to a first threshold, wherein the first prediction mode is an optical flow-based inter prediction mode; and derive gradient values in different directions based on prediction samples of the sub-block and an arithmetic shifting operation with a shift value S, and wherein S is an integer and S is not equal to 4.Join the waitlist — get patent alerts
Track US2021160511A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.