US2026065940A1PendingUtilityA1

Video generation method, apparatus, device, medium, product

Assignee: BEIJING ZITIAO NETWORK TECHNOLOGY CO LTDPriority: Aug 28, 2024Filed: Aug 22, 2025Published: Mar 5, 2026
Est. expiryAug 28, 2044(~18.1 yrs left)· nominal 20-yr term from priority
Inventors:HU QINGWEN
G06T 11/00G11B 27/031G06T 2207/30196G06T 2207/20081G06T 2207/10016G06T 2207/20084G06T 7/11G06T 7/55
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The disclosure discloses a video generation method, an apparatus, a device, a medium and a product. The method comprises: firstly, obtaining a first video so that at least one frame image in the first video includes a target object; then, for any frame image in the first video, performing body shape parameter prediction processing on the image to obtain the body shape parameter prediction result corresponding to the image, and performing adjustment processing on the body shape parameter prediction result corresponding to the image based on the body shape adjustment information specified for the target object to obtain the body shape parameter adjustment result corresponding to the image; finally, generating a second video based on the first video and the body shape parameter adjustment result corresponding to at least one image in the first video.

Claims

exact text as granted — not AI-modified
I/We claim: 
     
         1 . A video generation method, comprising:
 obtaining a first video, at least one frame image in the first video comprising a target object;   for the at least one frame image in the first video, obtaining a body shape parameter prediction result corresponding to the frame image by performing body shape parameter prediction processing on the frame image, the body shape parameter prediction result being used to describe a body shape of the target object in the frame image;   for any frame image in the first video, obtaining a body shape parameter adjustment result corresponding to the frame image by performing adjustment processing on the body shape parameter prediction result corresponding to the frame image according to body shape adjustment information specified for the target object; and   generating a second video according to the first video and a body shape parameter adjustment result corresponding to the at least one frame image in the first video, the i-th frame image in the second video being used to represent a body shape parameter adjustment result for the i-th frame image in the first video, the i-th frame image in the second video being generated according to a plurality of frame images arranged in sequence in the first video and body shape parameter adjustment results corresponding to the plurality of frame images, a distance between an arrangement position of at least one frame image of the plurality of frame images in the first video and an arrangement position of the i-th frame image in the first video being not greater than a preset distance threshold, i being a positive integer, and i being less than or equal to a number of image frames in the second video or a number of image frames in the first video.   
     
     
         2 . The method according to  claim 1 , wherein a body shape described by the i-th frame image in the second video is different from a body shape described by the i-th frame image in the first video, and other information except for the body shape described by the i-th frame image in the second video is consistent with information described by the i-th frame image in the first video except for the body shape; and/or,
 the body shape parameter prediction result corresponding to at least one frame image in the second video remains consistent.   
     
     
         3 . The method according to  claim 1 , wherein the second video is generated using a target model; and
 the target model comprises at least one processing unit, the processing unit comprises an image-level processing module and a time sequence module, the image-level processing module is configured to separately implement a separate processing procedure for the at least one frame image of the plurality of frame images, and the time sequence module is configured to perform attention processing on execution results of separate processing procedures for the plurality of frame images in a time direction.   
     
     
         4 . The method according to  claim 3 , wherein the at least one processing unit comprises a plurality of up-sampling modules, a time sequence module corresponding to at least one of the up-sampling module, a plurality of down-sampling modules, and a time sequence module corresponding to at least one of the down-sampling module;
 for any up-sampling module of the up-sampling modules, a time sequence module corresponding to the up-sampling module is configured to perform attention processing on output data of the up-sampling module in the time direction; and   for any down-sampling module of the down-sampling modules, a time sequence module corresponding to the down-sampling module is configured to perform attention processing on output data of the down-sampling module in the time direction.   
     
     
         5 . The method according to  claim 3 , wherein the attention processing is further implemented according to time information of the at least one frame image of the plurality of frame images; and
 for any frame image of the plurality of frame images, time information of the frame image is configured to describe an arrangement position of the frame image in the first video.   
     
     
         6 . The method according to  claim 3 , wherein the time sequence module comprises an integration network, a first conversion network, at least one self-attention network, a second conversion network, and a splitting network;
 the integration network is configured to perform integration processing on input data of the time sequence module to obtain an integration result, a size of the integration result being different from a size of the input data of the time sequence module, and the input data of the time sequence module comprising the execution results of separate processing procedures for the plurality of frame images;   the first conversion network is configured to convert the integration result from a first expression to a second expression to obtain a first conversion result, the first expression being used for describing an image feature, and the second expression being used for describing a video feature;   the at least one self-attention network is configured to perform self-attention processing on the first conversion result to obtain a self-attention processing result;   the second conversion network is configured to convert the self-attention processing result from the second expression to the first expression to obtain a second conversion result; and   the splitting network is configured to perform splitting processing on the second conversion result to obtain a splitting result, a size of the splitting result being the same as a size of the input data of the time sequence module.   
     
     
         7 . The method according to  claim 1 , wherein the second video is generated using a target model;
 the target model comprises at least one time sequence module; and   a training process of the target model comprises:   setting part or all parameters of the at least one time sequence module in the target model to zero, and training other parts of the target model except for the at least one time sequence module according to a first image and a second image, an object described by the first image being the same as an object described by the second image; and   freezing parameters of other parts of the target model except for the at least one time sequence module, and training the at least one time sequence module in the target model according to a first image sequence and a second image sequence, an object described by the first image sequence being the same as an object described by the second image sequence.   
     
     
         8 . The method according to  claim 7 , wherein the time sequence module comprises a second conversion network; and
 setting part or all parameters of the at least one time sequence module in the target model to zero comprises:   setting a parameter of the second conversion network of the at least one time sequence module in the target model to zero.   
     
     
         9 . The method according to  claim 7 , wherein a training process of the at least one time sequence module comprises:
 obtaining the first image sequence, the second image sequence, and label information corresponding to at least one frame image in the second image sequence;   obtaining a processed sequence by performing mask processing on at least one frame image in the first image sequence, the processed sequence comprising a mask processing result of the at least one frame image, and for any frame image of the at least one frame image, information described by a mask processing result of the frame image being less than information described by the frame image;   determining prediction information corresponding to the at least one frame image in the second image sequence according to the target model, the processed sequence and the body shape parameter prediction result corresponding to the at least one frame image in the second image sequence; and   updating the at least one time sequence module in the target model according to the prediction information corresponding to the at least one frame image in the second image sequence and the label information corresponding to the at least one frame image in the second image sequence.   
     
     
         10 . The method according to  claim 7 , wherein the training process of the at least one time sequence module comprises:
 obtaining label noise corresponding to at least one frame image in the second image sequence;   for any frame image in the second image sequence, obtaining a noise adding result corresponding to the frame image by performing noise adding processing on the frame image according to the label noise corresponding to the frame image;   determining prediction information corresponding to the at least one frame image in the second image sequence according to the target model, the first image sequence, the body shape parameter prediction result corresponding to the at least one frame image in the second image sequence, and a noise adding result corresponding to the at least one frame image in the second image sequence, the prediction information comprising predicted noise and a predicted image;   determining a first loss according to the label noise corresponding to the at least one frame image in the second image sequence and the predicted noise corresponding to the at least one frame image in the second image sequence;   determining a second loss according to a region segmentation result of the at least one frame image in the second image sequence, an object position detection result of the at least one frame image in the second image sequence, and a region segmentation result of a predicted image corresponding to the at least one frame image in the second image sequence; and   updating the at least one time sequence module in the target model according to a sum between the first loss and the second loss.   
     
     
         11 . An electronic device comprising: a processor and a memory;
 the memory, configured to store instructions or computer programs;   the processor, configured to execute the instructions or the computer programs in the memory, for causing the electronic device to:   obtain a first video, at least one frame image in the first video comprising a target object;   for any frame image in the first video, obtain a body shape parameter prediction result corresponding to the frame image by performing body shape parameter prediction processing on the frame image, the body shape parameter prediction result being used to describe a body shape of the target object in the frame image;   for any frame image in the first video, obtain a body shape parameter adjustment result corresponding to the frame image by performing adjustment processing on the body shape parameter prediction result corresponding to the frame image according to body shape adjustment information specified for the target object; and   generate a second video according to the first video and a body shape parameter adjustment result corresponding to the at least one frame image in the first video, the i-th frame image in the second video being used to represent a body shape parameter adjustment result for the i-th frame image in the first video, the i-th frame image in the second video being generated according to a plurality of frame images arranged in sequence in the first video and body shape parameter adjustment results corresponding to the plurality of frame images, a distance between an arrangement position of at least one frame image of the plurality of frame images in the first video and an arrangement position of the i-th frame image in the first video being not greater than a preset distance threshold, i being a positive integer, and i being less than or equal to a number of image frames in the second video or a number of image frames in the first video.   
     
     
         12 . The electronic device according to  claim 11 , wherein a body shape described by the i-th frame image in the second video is different from a body shape described by the i-th frame image in the first video, and other information except for the body shape described by the i-th frame image in the second video is consistent with information described by the i-th frame image in the first video except for the body shape; and/or,
 the body shape parameter prediction result corresponding to at least one frame image in the second video remains consistent.   
     
     
         13 . The electronic device according to  claim 11 , wherein the second video is generated using a target model; and
 the target model comprises at least one processing unit, the processing unit comprises an image-level processing module and a time sequence module, the image-level processing module is configured to separately implement a separate processing procedure for the at least one frame image of the plurality of frame images, and the time sequence module is configured to perform attention processing on execution results of separate processing procedures for the plurality of frame images in a time direction.   
     
     
         14 . The electronic device according to  claim 13 , wherein the at least one processing unit comprises a plurality of up-sampling modules, a time sequence module corresponding to at least one of the up-sampling module, a plurality of down-sampling modules, and a time sequence module corresponding to at least one of the down-sampling module;
 for any up-sampling module of the up-sampling modules, a time sequence module corresponding to the up-sampling module is configured to perform attention processing on output data of the up-sampling module in the time direction; and   for any down-sampling module of the down-sampling modules, a time sequence module corresponding to the down-sampling module is configured to perform attention processing on output data of the down-sampling module in the time direction.   
     
     
         15 . The electronic device according to  claim 13 , wherein the attention processing is further implemented according to time information of the at least one frame image of the plurality of frame images; and
 for the at least one frame image of the plurality of frame images, time information of the frame image is configured to describe an arrangement position of the frame image in the first video.   
     
     
         16 . The electronic device according to  claim 13 , wherein the time sequence module comprises an integration network, a first conversion network, at least one self-attention network, a second conversion network, and a splitting network;
 the integration network is configured to perform integration processing on input data of the time sequence module to obtain an integration result, a size of the integration result being different from a size of the input data of the time sequence module, and the input data of the time sequence module comprising the execution results of separate processing procedures for the plurality of frame images;   the first conversion network is configured to convert the integration result from a first expression to a second expression to obtain a first conversion result, the first expression being used for describing an image feature, and the second expression being used for describing a video feature;   the at least one self-attention network is configured to perform self-attention processing on the first conversion result to obtain a self-attention processing result;   the second conversion network is configured to convert the self-attention processing result from the second expression to the first expression to obtain a second conversion result; and   the splitting network is configured to perform splitting processing on the second conversion result to obtain a splitting result, a size of the splitting result being the same as a size of the input data of the time sequence module.   
     
     
         17 . The electronic device according to  claim 11 , wherein the second video is generated using a target model;
 the target model comprises at least one time sequence module; and   the instructions or the computer programs causing the electronic device to perform a training process of the target model comprise instructions to:   set part or all parameters of the at least one time sequence module in the target model to zero, and train other parts of the target model except for the at least one time sequence module according to a first image and a second image, an object described by the first image being the same as an object described by the second image; and   freeze parameters of other parts of the target model except for the at least one time sequence module, and train the at least one time sequence module in the target model according to a first image sequence and a second image sequence, an object described by the first image sequence being the same as an object described by the second image sequence.   
     
     
         18 . The electronic device according to  claim 17 , wherein the time sequence module comprises a second conversion network; and
 the instructions or the computer programs causing the electronic device to set part or all parameters of the at least one time sequence module in the target model to zero comprise instructions to:   set a parameter of the second conversion network of the at least one time sequence module in the target model to zero.   
     
     
         19 . The electronic device according to  claim 17 , wherein the instructions or the computer programs causing the electronic device to perform a training process of the at least one time sequence module comprise instructions to:
 obtain the first image sequence, the second image sequence, and label information corresponding to at least one frame image in the second image sequence;   obtain a processed sequence by performing mask processing on at least one frame image in the first image sequence, the processed sequence comprising a mask processing result of the at least one frame image, and for any frame image of the at least one frame image, information described by a mask processing result of the frame image being less than information described by the frame image;   determine prediction information corresponding to the at least one frame image in the second image sequence according to the target model, the processed sequence and the body shape parameter prediction result corresponding to the at least one frame image in the second image sequence; and   update the at least one time sequence module in the target model according to the prediction information corresponding to the at least one frame image in the second image sequence and the label information corresponding to the at least one frame image in the second image sequence.   
     
     
         20 . A non-transitory computer-readable medium, wherein instructions or a computer program are stored in the computer-readable medium, and when the instructions or the computer program are run on a device, causing the device to:
 obtain a first video, at least one frame image in the first video comprising a target object;   for any frame image in the first video, obtain a body shape parameter prediction result corresponding to the frame image by performing body shape parameter prediction processing on the frame image, the body shape parameter prediction result being used to describe a body shape of the target object in the frame image;   for any frame image in the first video, obtain a body shape parameter adjustment result corresponding to the frame image by performing adjustment processing on the body shape parameter prediction result corresponding to the frame image according to body shape adjustment information specified for the target object; and   generate a second video according to the first video and a body shape parameter adjustment result corresponding to the at least one frame image in the first video, the i-th frame image in the second video being used to represent a body shape parameter adjustment result for the i-th frame image in the first video, the i-th frame image in the second video being generated according to a plurality of frame images arranged in sequence in the first video and body shape parameter adjustment results corresponding to the plurality of frame images, a distance between an arrangement position of at least one frame image of the plurality of frame images in the first video and an arrangement position of the i-th frame image in the first video being not greater than a preset distance threshold, i being a positive integer, and i being less than or equal to a number of image frames in the second video or a number of image frames in the first video.

Join the waitlist — get patent alerts

Track US2026065940A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.