US11876993B2ActiveUtilityA1

Signaling of combined intra-inter prediction

Assignee: BEIJING BYTEDANCE NETWORK TECH CO LTDPriority: Mar 21, 2019Filed: Jun 23, 2021Granted: Jan 16, 2024
Est. expiryMar 21, 2039(~12.7 yrs left)· nominal 20-yr term from priority
H04N 19/46H04N 19/105H04N 19/107H04N 19/126H04N 19/159H04N 19/176H04N 19/521H04N 19/58H04N 19/96H04N 19/503H04N 19/593H04N 19/186H04N 19/70H04N 19/52
83
PatentIndex Score
1
Cited by
86
References
20
Claims

Abstract

The present application relates to signaling of combined intra-inter prediction. A method for processing video includes: coding, during a conversion between a current video block in a video data and a bitstream representation of the current video block, a combined inter-intra prediction (CIIP) flag for the current video block by a context model based coding without referring to a CIIP flag of one or more neighboring video blocks to the current video block, and performing, at least by applying the combined inter-intra prediction (CIIP) flag of the current video block, the conversion.

Claims

exact text as granted — not AI-modified
The invention claimed is: 
     
       1. A method for processing video, comprising:
 determining, during a conversion between a current video block in a video data and a bitstream, two candidates and a combined inter-intra prediction flag for the current video block by a context model-based coding without referring to a combined inter-intra prediction flag of one or more neighboring video blocks to the current video block,
 wherein the combined inter-intra prediction flag is used to indicate whether a combined inter-intra prediction mode is used, and 
 wherein, in the combined inter-intra prediction mode, a prediction signal of the current video block is generated at least based on an intra prediction signal and an inter prediction signal; 
 
 comparing a first information of the two candidates to determine whether to add at least one of the two candidates to a candidate list constructed for the current video block; and 
 performing, by at least applying the combined inter-intra prediction flag of the current video block, the conversion based on a result of the comparing, 
 wherein the first information of two candidates excludes a second information related to a coding mode, and 
 wherein the second information related to the coding mode comprises at least one of:
 a flag of the coding mode, or 
 an intra mode that is used in the coding mode, 
 
 wherein a weight pair is determined according to a prediction mode of the one or more neighboring video blocks of the current video block, 
 wherein, when a neighboring video block is coded with the combined inter-intra prediction mode, the neighboring video block is treated as a block coded with an inter prediction mode which is a non-intra prediction mode, 
 wherein the one or more neighboring video blocks comprise a video block covering a location (xCb−1, yCb−1+(cbHeight<<a)) and a video block covering a location (xCb−1+(cbWidth<<a), yCb−1), wherein (xCb, yCb) is a location of a top-left sample of the current video block, cbWidth and cbHeight are a width and a height of the current video block, respectively, and a is determined using a cIdx of the current video block, and wherein cIdx is a variable specifying a color component index for the current video block. 
 
     
     
       2. The method of  claim 1 , wherein in response to the current video block being coded with the combined inter-intra prediction mode, intra-predication mode information of the current video block is not stored. 
     
     
       3. The method of  claim 1 , wherein in response to the current video block being coded with the combined inter-intra prediction mode, the current video block is deemed as a video block coded with an inter-prediction mode. 
     
     
       4. The method of  claim 1 , wherein the conversion comprises decoding the current video block from the bitstream. 
     
     
       5. The method of  claim 1 , wherein the conversion comprises encoding the current video block into the bitstream. 
     
     
       6. The method of  claim 1 , wherein the two candidates comprise two spatial merge candidates. 
     
     
       7. The method of  claim 1 , wherein the two candidates comprise a spatial merge candidate in the candidate list and a candidate from a history-based motion vector prediction table. 
     
     
       8. The method of  claim 1 ,
 wherein the weight pair comprising a first weight for a first prediction result of the current video block and a second weight for a second prediction result of the current video block, based on the one or more neighboring video blocks of the current video block, wherein the first prediction result is generated by an intra prediction mode, and the second prediction result is generated by an inter prediction mode; and 
 determining a prediction result of the current video block based on a weighted sum of the first prediction result and the second prediction result, 
 wherein the weight pair is determined based on two or more neighboring video blocks 
 wherein,
 when all of the two or more neighboring video blocks are coded with the intra prediction mode, the weight pair is a first candidate weight pair, 
 when all of the two or more neighboring video blocks are coded with a non-intra prediction mode, the weight pair is a second candidate weight pair different from the first candidate weight pair, and 
 otherwise, the weight pair is a third candidate weight pair different from the first candidate weight pair and the second candidate weight pair. 
 
 
     
     
       9. The method of  claim 8 , wherein the first candidate weight pair is (3, 1), the second candidate weight is (1, 3) and the third candidate weight pair is (2, 2), and wherein for (x, y), x is the first weight and y is the second weight. 
     
     
       10. The method of  claim 1 , further comprising:
 determining, in response to the current video block being coded with the combined inter-intra prediction mode, a weight pair comprising a first weight for a first prediction result of the current video block and a second weight for a second prediction result of the current video block, based on the one or more neighboring video blocks of the current video block, wherein the first prediction result is generated by an intra prediction mode, and the second prediction result is generated by an inter prediction mode; and 
 determining a prediction result of the current video block based on a weighted sum of the first prediction result and the second prediction result, 
 wherein the prediction result is obtained by applying the weight pair for the intra prediction result and the inter prediction result as:
     P _=( w Inter* P _inter+ w Intra* P _intra+offset)>> N , and 
 
 wherein P_ is the prediction result, P_inter is the first prediction result, P_intra is the second prediction result, (winter, wintra) is the weight pair, and each of offset and N is an integer. 
 
     
     
       11. The method of  claim 1 ,
 wherein a fixed context is determined without referring to whether the combined inter-intra prediction mode is used for one or more neighboring video blocks to the current video block and is used in the context model-based coding of the combined inter-intra prediction flag of the current video block. 
 
     
     
       12. The method of  claim 1 , further comprising:
 in response to the current video block being coded with the combined inter-intra prediction mode, directly using a planar mode to generate the intra prediction signal and excluding checking the combined inter-intra prediction flag of one or more neighboring video blocks to the current video block, wherein the planar mode is the only intra mode used for a video block coded with the combined inter-intra prediction mode. 
 
     
     
       13. The method of  claim 12 , wherein the planar mode is used for an intra prediction mode determination process of subsequent coded video blocks. 
     
     
       14. The method of  claim 13 , wherein during a conversion of a second video block which is one of subsequent coded video blocks of the current video block, the planar mode is added to an intra prediction mode candidate list of the second video block. 
     
     
       15. The method of  claim 14 , wherein the intra prediction mode candidate list includes a most-probably-mode candidate list. 
     
     
       16. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:
 determine, during a conversion between a current video block in a video data and a bitstream of the current video block, two candidates and a combined inter-intra prediction flag for the current video block by a context model-based coding without referring to a combined inter-intra prediction flag of one or more neighboring video blocks to the current video block,
 wherein the combined inter-intra prediction flag is used to indicate whether a combined inter-intra prediction mode is used, and 
 wherein, in the combined inter-intra prediction mode, a prediction signal of the current video block is generated at least based on an intra prediction signal and an inter prediction signal; 
 
 compare a first information of the two candidates to determine whether to add at least one of the two candidates to a candidate list constructed for the current video block; and 
 perform, by at least applying the combined inter-intra prediction flag of the current video block, the conversion based on a result of comparing the first information, 
 wherein the first information of two candidates excludes a second information related to a coding mode, and 
 wherein the second information related to the coding mode comprises at least one of:
 a flag of the coding mode, or 
 an intra mode that is used in the coding mode, 
 
 wherein a weight pair is determined according to a prediction mode of the one or more neighboring video blocks of the current video block, 
 wherein, when a neighboring video block is coded with the combined inter-intra prediction mode, the neighboring video block is treated as a block coded with an inter prediction mode which is a non-intra prediction mode, 
 wherein the one or more neighboring video blocks comprise a video block covering a location (xCb−1, yCb−1+(cbHeight<<a)) and a video block covering a location (xCb−1+(cbWidth<<a), yCb−1), wherein (xCb, yCb) is a location of a top-left sample of the current video block, cbWidth and cbHeight are a width and a height of the current video block, respectively, and a is determined using a cIdx of the current video block, and wherein cIdx is a variable specifying a color component index for the current video block. 
 
     
     
       17. The apparatus of  claim 16 , wherein the instructions further cause the processor to:
 in response to the current video block being coded with the combined inter-intra prediction mode, directly use a planar mode to generate the intra prediction signal and exclude checking the combined inter-intra prediction flag of one or more neighboring video blocks to the current video block, 
 wherein the planar mode is the only intra mode used for a video block coded with the combined inter-intra prediction mode, and 
 wherein the planar mode is used for an intra prediction mode determination process of subsequent coded video blocks. 
 
     
     
       18. The apparatus of  claim 17 , wherein during a conversion of a second video block which is one of subsequent coded video blocks of the current video block, the planar mode is added to an intra prediction mode candidate list of the second video block. 
     
     
       19. A non-transitory computer-readable storage medium storing instructions that cause a processor to:
 determine, during a conversion between a current video block in a video data and a bitstream, two candidates and a combined inter-intra prediction flag for the current video block by a context model-based coding without referring to a combined inter-intra prediction flag of one or more neighboring video blocks to the current video block,
 wherein the combined inter-intra prediction flag is used to indicate whether a combined inter-intra prediction mode is used, and 
 wherein, in the combined inter-intra prediction mode, a prediction signal of the current video block is generated at least based on an intra prediction signal and an inter prediction signal; 
 
 compare a first information of the two candidates to determine whether to add at least one of the two candidates to a candidate list constructed for the current video block; and 
 perform, by at least applying the combined inter-intra prediction flag of the current video block, the conversion based on a result of comparing the first information, 
 wherein the first information of two candidates excludes a second information related to a coding mode, and 
 wherein the second information related to the coding mode comprises at least one of:
 a flag of the coding mode, or 
 an intra mode that is used in the coding mode, 
 
 wherein a weight pair is determined according to a prediction mode of the one or more neighboring video blocks of the current video block, 
 wherein, when a neighboring video block is coded with the combined inter-intra prediction mode, the neighboring video block is treated as a block coded with an inter prediction mode which is a non-intra prediction mode, 
 wherein the one or more neighboring video blocks comprise a video block covering a location (xCb−1, yCb−1+(cbHeight<<a)) and a video block covering a location (xCb−1+(cbWidth<<a), yCb−1), wherein (xCb, yCb) is a location of a top-left sample of the current video block, cbWidth and cbHeight are a width and a height of the current video block, respectively, and a is determined using a cIdx of the current video block, and wherein cIdx is a variable specifying a color component index for the current video block. 
 
     
     
       20. A non-transitory computer-readable recording medium storing a bitstream of a video data which is generated by a method performed by a video processing apparatus, wherein the method comprises:
 determining, for a current video block in a video data, two candidates and a combined inter-intra prediction flag for the current video block by a context model-based coding without referring to a combined inter-intra prediction flag of one or more neighboring video blocks to the current video block,
 wherein the combined inter-intra prediction flag is used to indicate whether a combined inter-intra prediction mode is used, and 
 wherein, in the combined inter-intra prediction mode, a prediction signal of the current video block is generated at least based on an intra prediction signal and an inter prediction signal; 
 
 comparing a first information of the two candidates to determine whether to add at least one of the two candidates to a candidate list constructed for the current video block; and 
 generating, by at least applying the combined inter-intra prediction flag of the current video block, the bitstream based on a result of the comparing, 
 wherein the first information of two candidates excludes a second information related to a coding mode, and 
 wherein the second information related to the coding mode comprises at least one of:
 a flag of the coding mode, or 
 an intra mode that is used in the coding mode, 
 wherein a weight pair is determined according to a prediction mode of the one or more neighboring video blocks of the current video block, 
 
 wherein, when a neighboring video block is coded with the combined inter-intra prediction mode, the neighboring video block is treated as a block coded with an inter prediction mode which is a non-intra prediction mode, 
 wherein the one or more neighboring video blocks comprise a video block covering a location (xCb−1, yCb−1+(cbHeight<<a)) and a video block covering a location (xCb−1+(cbWidth<<a), yCb−1), wherein (xCb, yCb) is a location of a top-left sample of the current video block, cbWidth and cbHeight are a width and a height of the current video block, respectively, and a is determined using a cIdx of the current video block, and wherein cIdx is a variable specifying a color component index for the current video block.

Join the waitlist — get patent alerts

Track US11876993B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.