Affine motion prediction-based video decoding method and device using subblock-based temporal merge candidate in video coding system
Abstract
A video decoding method performed by a decoding device according to the present document is characterized by including: a step for deriving reference subblocks in a reference picture on the basis of the motion vector of an adjacent block on the left side of the current block; a step for deriving a subblock-based temporal merge candidate for the current block on the basis of motion information about the reference subblocks; a step for forming an affine merge candidate list for the current block, the affine merge candidate list including the subblock-based temporal merge candidate; a step for deriving motion information about subblocks of the current block on the basis of the affine merge candidate list; a step for deriving prediction samples for the current block on the basis of the motion information about the subblocks; and a step for generating a reconstructed picture on the basis of the prediction samples.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An image decoding apparatus comprising:
a memory; and at least one processor connected to the memory, the at least one processor configured to: derive a reference block in a reference picture based on a left neighboring block of a current block; derive a subblock-based temporal merging candidate for the current block based on the reference block; derive an inherited affine candidate for the current block; derive a constructed affine candidate for the current block; construct a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate; derive motion information of sub-blocks of the current block based on the subblock merge candidate list; derive prediction samples for the current block based on motion information of the sub-blocks; generate a reconstructed picture based on the prediction samples; and filter the reconstructed picture based on at least one of a deblocking filter, a sample adaptive offset, or adaptive loop filter, wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a v component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.
2 . The image decoding apparatus of claim 1 , wherein a position of the reference block is derived based on a motion vector of the left neighboring block.
3 . The image decoding apparatus of claim 2 , wherein the motion vector for deriving the position of the reference block is fixed to the motion vector of the left neighboring block.
4 . The image decoding apparatus of claim 1 ; wherein the at least one processor, to derive the motion information of the sub-blocks of the current block, further configured to:
select the subblock-based temporal merging candidate from the subblock merge candidate list; and derive the motion information of the sub-blocks of the current block based on the subblock-based temporal merging candidate.
5 . The image decoding apparatus of claim 4 , wherein motion information of a target sub-block among the sub-blocks is derived based on motion information of a collocated sub-block for the target sub-block included in the subblock-based temporal merging candidate.
6 . An image encoding apparatus comprising:
a memory; and at least one processor connected to the memory, the at least one processor configured to: derive a reference block in a reference picture based on a left neighboring block of a current block; derive a subblock-based temporal merging candidate for the current block based on motion information of the reference block; derive an inherited affine candidate for the current block; derive a constructed affine candidate for the current block; construct a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate; derive motion information of sub-blocks of the current block based on the subblock merge candidate list; derive prediction samples for the current block based on the motion information of the sub-blocks; determine to apply at least one of a deblocking filter, a sample adaptive offset, or adaptive loop filter; and encode image information including prediction information for the current block, wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a y component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.
7 . The image encoding apparatus of claim 6 , wherein a position of the reference block is derived based on a motion vector of the left neighboring block.
8 . The image encoding apparatus of claim 7 , wherein the motion vector for deriving the position of the reference block is fixed to the motion vector of the left neighboring block.
9 . The image encoding apparatus of claim 6 , wherein the at least one processor, to derive the motion information of the sub-blocks of the current block, further configured to:
select the subblock-based temporal merging candidate from the subblock merge candidate list; and derive the motion information of the sub-blocks of the current block based on the subblock-based temporal merging candidate.
10 . The image encoding apparatus of claim 9 , wherein motion information of a target sub-block among the sub-blocks is derived based on motion information of a collocated sub-block for the target sub-block included in the subblock-based temporal merging candidate.
11 . A apparatus for transmitting a bitstream for an image, the apparatus comprising:
at least one processor configured to obtain the bitstream generated by an image encoding apparatus; and a transmitter configured to transmit the bitstream, the bitstream is generated by: deriving a reference block in a reference picture based on a left neighboring block of a current block; deriving a subblock-based temporal merging candidate for the current block based on motion information of the reference block; deriving an inherited affine candidate for the current block; deriving a constructed affine candidate for the current block; constructing a subblock merge candidate list for the current block including the subblock-based temporal merging candidate, the inherited affine candidate and the constructed affine candidate; deriving motion information of sub-blocks of the current block based on the subblock merge candidate list; deriving prediction samples for the current block based on the motion information of the sub-blocks; determining to apply at least one of a deblocking filter, a sample adaptive offset, or adaptive loop filter; and encoding image information including prediction information for the current block, wherein based on a size of the current block being W×H, and an x component of a top-left sample position of the current block being a and a y component of the top-left sample position being b, the left neighboring block is a block including a sample at (a−1, b+H−1) coordinates.Join the waitlist — get patent alerts
Track US2026019567A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.