Processing framework for temporal-consistent face manipulation in videos
Abstract
Embodiments are disclosed for generating temporally consistent manipulated videos. A method of generating temporally consistent manipulated videos comprises receiving a target appearance and an input digital video including a plurality of frames, generating a plurality of target appearance frames from the plurality of frames, training a video prediction network to generate a digital video wherein a subject of the digital video has its appearance modified to match the target appearance, providing the input digital video to the video prediction network, and generating, by the video prediction network, an output digital video wherein the subject of the output digital video has its appearance modified to match the target appearance.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
receiving a target appearance and an input digital video including a plurality of frames; generating a plurality of target appearance frames from the plurality of frames; training a video prediction network to generate a digital video wherein a subject of the digital video has its appearance modified to match the target appearance; providing the input digital video to the video prediction network; and generating, by the video prediction network, an output digital video wherein the subject of the output digital video has its appearance modified to match the target appearance.
2 . The method of claim 1 , wherein the plurality of frames represents a subset of the input digital video.
3 . The method of claim 1 , wherein generating a plurality of target appearance frames from the plurality of frames, further comprises:
processing the plurality of frames by a subject cropping network to generate a plurality of input cropped images.
4 . The method of claim 3 , further comprising:
providing the plurality of input cropped images to a subject manipulation network; identifying a plurality of latent representations corresponding to the target appearance in the plurality of input cropped images; and generating a plurality of target appearance cropped images using the plurality of latent representations.
5 . The method of claim 4 , further comprising:
blending the plurality of target appearance cropped images and the plurality of frames using a subject blending network to generate the plurality of target appearance frames.
6 . The method of claim 1 , wherein training a video prediction network to generate a digital video wherein a subject of the digital video has its appearance modified to match the target appearance, further comprises:
training a plurality of video manipulation networks, wherein the plurality of video manipulation networks include an encoder network, a first decoder network, and a second decoder network.
7 . The method of claim 6 , wherein training a plurality of video manipulation networks, wherein the plurality of video manipulation networks include an encoder network, a first decoder network, and a second decoder network, further comprises:
providing the plurality of frames and the plurality of target appearance frames to the encoder network; generating, by the encoder network, a representation of the plurality of frames and a representation of the plurality of target appearance frames; reconstructing, by the first decoder network, a plurality of reconstructed frames from the representation of the plurality of frames; reconstructing, by the second decoder network, a plurality of reconstructed target appearance frames from the representation of the plurality of target appearance frames; and training the first decoder network, second decoder network, and encoder network, by comparing the plurality of reconstructed frames to the plurality of frames and the plurality of reconstructed target appearance frames to the plurality of target appearance frames using a loss function.
8 . The method of claim 6 , wherein the video prediction network comprises the encoder network and the second decoder network.
9 . The method of claim 1 , wherein a subject of the input digital video includes a representation of a person's face and wherein the target appearance includes a change to an expression or appearance of the person's face.
10 . A non-transitory computer-readable storage medium including instructions stored thereon which, when executed by at least one processor, cause the at least one processor to:
receive a target appearance and an input digital video including a plurality of frames; generate a plurality of target appearance frames from the plurality of frames; train a video prediction network to generate a digital video wherein a subject of the digital video has its appearance modified to match the target appearance; provide the input digital video to the video prediction network; and generate, by the video prediction network, an output digital video wherein the subject of the output digital video has its appearance modified to match the target appearance.
11 . The non-transitory computer-readable storage medium of claim 10 , wherein the plurality of frames represents a subset of the input digital video.
12 . The non-transitory computer-readable storage medium of claim 10 , wherein to generate a plurality of target appearance frames from the plurality of frames, the instructions, when executed, further cause the at least one processor to:
process the plurality of frames by a subject cropping network to generate a plurality of input cropped images.
13 . The non-transitory computer-readable storage medium of claim 12 , wherein the instructions, when executed, further cause the at least one processor to:
provide the plurality of input cropped images to a subject manipulation network; identify a plurality of latent representations corresponding to the target appearance in the plurality of input cropped images; and generate a plurality of target appearance cropped images using the plurality of latent representations.
14 . The non-transitory computer-readable storage medium of claim 13 , wherein the instructions, when executed, further cause the at least one processor to:
blend the plurality of target appearance cropped images and the plurality of frames using a subject blending network to generate the plurality of target appearance frames.
15 . The non-transitory computer-readable storage medium of claim 10 , wherein to train a video prediction network to generate a digital video wherein a subject of the digital video has its appearance modified to match the target appearance, the instructions, when executed, further cause the at least one processor to:
train a plurality of video manipulation networks, wherein the plurality of video manipulation networks include an encoder network, a first decoder network, and a second decoder network.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein to train a plurality of video manipulation networks, wherein the plurality of video manipulation networks include an encoder network, a first decoder network, and a second decoder network, the instructions, when executed, further cause the at least one processor to:
provide the plurality of frames and the plurality of target appearance frames to the encoder network; generate, by the encoder network, a representation of the plurality of frames and a representation of the plurality of target appearance frames; reconstruct, by the first decoder network, a plurality of reconstructed frames from the representation of the plurality of frames; reconstruct, by the second decoder network, a plurality of reconstructed target appearance frames from the representation of the plurality of target appearance frames; and train the first decoder network, second decoder network, and encoder network, by comparing the plurality of reconstructed frames to the plurality of frames and the plurality of reconstructed target appearance frames to the plurality of target appearance frames using a loss function.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the video prediction network comprises the encoder network and the second decoder network.
18 . The non-transitory computer-readable storage medium of claim 10 , wherein a subject of the input digital video includes a representation of a person's face and wherein the target appearance includes a change to an expression or appearance of the person's face.
19 . A system comprising:
a memory component; and a processing device coupled to the memory component, the processing device to execute instructions stored on the memory component which cause the system to perform operations comprising:
generating a plurality of target appearance frames corresponding to an input digital video based on an appearance input defining a target appearance of a subject of the input digital video;
training a video prediction network using the plurality of target appearance frames and the input digital video;
providing the input digital video to the video prediction network; and
generating, by the video prediction network, a temporally consistent output digital video wherein the subject of the temporally consistent output digital video has its appearance modified to match the target appearance.
20 . The system of claim 19 , wherein the plurality of frames represents a subset of the input digital video.Join the waitlist — get patent alerts
Track US2023377339A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.