Method for processing image on basis of inter-prediction mode and device therefor
Abstract
Disclosed herein are a method for decoding a video signal and a device therefor. Specifically, a method for decoding an image based on an inter prediction mode may include: if a merge mode is applied to a current block, generating a merge candidate list by using a spatial neighboring block and a temporal neighboring block of the current block; obtaining a merge index indicating a candidate to be used for an inter prediction of the current block in the merge candidate list; deriving a motion vector of each of subblocks included in the current block based on a motion vector of the candidate used for the inter prediction of the current block; and generating a prediction sample of the current block by using the motion vector of each of subblocks.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1. A method of decoding a video signal, comprising:
generating a merge candidate list including at least one of a spatial neighboring block and a temporal neighboring block of a current block, and including candidates equal to or less than a maximum number of merge candidates, based on that a merge mode is applied to the current block;
obtaining a merge index indicating a candidate to be used for an inter prediction of the current block in the merge candidate list;
deriving motion vectors of each subblocks included in the current block based on motion vectors of the candidate indicated by the merge index;
generating a prediction block of the current block based on the motion vectors of each subblocks;
performing inverse-transform on transformed coefficients of the current block to generate a residual block of the current block; and
generating a reconstructed block of the current block based on the prediction block and the residual block,
wherein the step of generating the merge candidate list comprises:
based on that a width and a height of the current block are respectively greater than or equal to a pre-defined minimum size, adding the temporal neighboring block which is located in a collocated picture of the current block and specified by a motion vector of the spatial neighboring block to the merge candidate list,
wherein based on that the width or the height of the current block is smaller than the pre-defined minimum size, the temporal neighboring block is not included in the merge candidate list, and
wherein based on that the merge index indicates the temporal neighboring block as the candidate, the motion vectors of the each subblocks included in the current block are derived based on motion vectors of subblocks included in the temporal neighboring block.
2. The method of claim 1 , wherein the temporal neighboring block is specified by a motion vector of a first spatial neighboring block of the merge candidate list, the first spatial neighboring block being determined based on a pre-defined order among two or more spatial neighboring blocks of the current block.
3. The method of claim 1 , wherein the pre-defined minimum size is set to 8.
4. A method of encoding a video signal, comprising:
generating a merge candidate list including at least one of a spatial neighboring block and a temporal neighboring block of a current block, and including candidates equal to or less than a maximum number of merge candidates, based on that a merge mode is applied to the current block;
deriving motion vectors of each subblocks included in the current block based on motion vectors of a candidate included in the merge candidate list;
generating a prediction block of the current block based on the motion vectors of each subblocks;
generating a residual block of the current block based on the prediction block; and
generating a merge index indicating the candidate in the merge candidate list;
performing transform on the residual block,
wherein the step of generating the merge candidate list comprises:
based on that a width and a height of the current block being respectively greater than or equal to a pre-defined minimum size, adding the temporal neighboring block which is located in a collocated picture of the current block and specified by a motion vector of the spatial neighboring block to the merge candidate list,
wherein based on that the width or the height of the current block is smaller than the pre-defined minimum size, the temporal neighboring block is not included in the merge candidate list, and
wherein based on the motion vector of the temporal neighboring block in the merge candidate list being used for deriving the motion vectors of the each subblocks of the current block, the motion vectors of the each subblocks of the current block are derived based on motion vectors of subblocks divided from the temporal neighboring block.
5. The method of claim 4 , wherein the temporal neighboring block is specified by a motion vector of a first spatial neighboring block of the merge candidate list, the first spatial neighboring block being determined based on a pre-defined order among two or more spatial neighboring blocks of the current block.
6. The method of claim 4 , wherein the pre-defined minimum size is set to 8.
7. A transmission method for data comprising a bitstream for an image, the method comprising:
obtaining the bitstream for the image; and
transmitting the data comprising the bitstream,
wherein the bitstream is generated by performing the steps of:
generating a merge candidate list including at least one of a spatial neighboring block and a temporal neighboring block of a current block, and including candidates equal to or less than a maximum number of merge candidates, based on that a merge mode is applied to the current block;
deriving motion vectors of each subblocks included in the current block based on motion vectors of a candidate included in the merge candidate list;
generating a prediction block of the current block based on the motion vectors of each subblocks;
generating a residual block of the current block based on the prediction block; and
generating a merge index indicating the candidate in the merge candidate list;
performing transform on the residual block,
wherein the step of generating the merge candidate list comprises:
based on that a width and a height of the current block being respectively greater than or equal to a pre-defined minimum size, adding the temporal neighboring block which is located in a collocated picture of the current block and specified by a motion vector of the spatial neighboring block to the merge candidate list,
wherein based on that the width or the height of the current block is smaller than the pre-defined minimum size, the temporal neighboring block is not included in the merge candidate list, and
wherein based on the motion vector of the temporal neighboring block in the merge candidate list being used for deriving the motion vectors of the each subblocks of the current block, the motion vectors of the each subblocks of the current block are derived based on motion vectors of subblocks divided from the temporal neighboring block.Join the waitlist — get patent alerts
Track US12096004B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.