Applying and blending new textures to surfaces across frames of a video sequence
Abstract
Embodiments are disclosed for a process of applying and blending new textures to surfaces across frames of a video sequence. The method may include obtaining a new texture for a selected region of a video frame of a video sequence. The method may further comprise generating a mesh for the selected region of the first video frame that includes a plurality of control points. The method may further comprise determining control point location data for each of the plurality of control points for additional video frames of the video sequence and using the control point location data to generate a plurality of warped video frames by applying the new texture to the additional video. The method may further comprise generating blended video frames by blending the new texture in the warped video frames and providing a modified version of the video sequence using the generated blended video frames.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
obtaining a new texture for a selected region of a first video frame of a video sequence; generating a mesh for the selected region of the first video frame of the video sequence, wherein the mesh includes a plurality of control points arranged within the mesh; determining control point location data for each of the plurality of control points for additional video frames of the video sequence; generating warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence; generating blended video frames by blending the new texture in the warped video frames; and providing a modified version of the video sequence using the generated blended video frames.
2 . The method of claim 1 , wherein determining the control point location data for each of the plurality of control points for the additional video frames of the video sequence further comprises:
for each pair of consecutive video frames of the video sequence:
determining motions of each of the plurality of control points from the first video frame of the video sequence to a second video frame of the video sequence from an optical flow mapping of the video sequence,
generating a transformation function representing deformation of the plurality of control points using the determined motions of each of the plurality of control points,
generating locations for each of the plurality of control points in the second video frame of the video sequence by warping the plurality of control points from the first video frame of the video sequence using the generated transformation function, and
storing the generated locations for each of the plurality of control points in the second video frame of the video sequence as the control point location data.
3 . The method of claim 2 , further comprising:
determining a reverse optical flow location of a control point in the first video frame of the video sequence using a reverse optical flow mapping from the second video frame of the video sequence to the first video frame of the video sequence; calculating a distance between an original location of the control point and a reverse optical flow location of the control point; and removing the control point from the plurality of control points when the calculated distance is greater than a threshold value.
4 . The method of claim 2 , wherein generating the warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence further comprises:
determining first control point location data for the first video frame and second control point location data for a target video frame of the video sequence from the generated control point location data; generating a warping function using the first control point location data and the second control point location data; and warping the new texture from the first video frame to the target video frame using the generated warping function.
5 . The method of claim 4 , wherein generating the blended video frames by blending the new texture in the warped video frames further comprises:
for each video frame of the additional video frames of the video sequence:
generating a preliminary blend by blending the new texture with the region in the video frame,
computing high-frequency residuals lost during the generating of the preliminary blending by blending a transparent image with the video frame, and
computing a final blend of the new texture with the region in the video frame by merging the preliminary blend and the computed high-frequency residuals.
6 . The method of claim 2 , wherein generating the transformation function representing the deformation of the plurality of control points using the determined motions of each of the plurality of control points further comprises:
generating a first transformation function representing rigid deformation and a second transformation function representing non-rigid deformation; and generating the transformation function by applying a weighting to the first transformation function and the second transformation function.
7 . The method of claim 1 , wherein generating the blended video frames by blending the new texture in the warped video frames comprises:
receiving a second input including a prompt, the prompt indicating a requested texture for the selected region of the first video frame.
8 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:
obtaining a new texture for a selected region of a first video frame of a video sequence; generating a mesh for the selected region of the first video frame of the video sequence, wherein the mesh includes a plurality of control points arranged within the mesh; determining control point location data for each of the plurality of control points for additional video frames of the video sequence; generating warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence; generating blended video frames by blending the new texture in the warped video frames; and providing a modified version of the video sequence using the generated blended video frames.
9 . The non-transitory computer-readable medium of claim 8 , wherein the instructions to determine the control point location data for each of the plurality of control points for the additional video frames of the video sequence further comprises:
for each pair of consecutive video frames of the video sequence:
determining motions of each of the plurality of control points from the first video frame of the video sequence to a second video frame of the video sequence from an optical flow mapping of the video sequence,
generating a transformation function representing deformation of the plurality of control points using the determined motions of each of the plurality of control points,
generating locations for each of the plurality of control points in the second video frame of the video sequence by warping the plurality of control points from the first video frame of the video sequence using the generated transformation function, and
storing the generated locations for each of the plurality of control points in the second video frame of the video sequence as the control point location data.
10 . The non-transitory computer-readable medium of claim 9 , storing instructions that further cause the processing device to perform operations comprising:
determining a reverse optical flow location of a control point in the first video frame of the video sequence using a reverse optical flow mapping from the second video frame of the video sequence to the first video frame of the video sequence; calculating a distance between an original location of the control point and a reverse optical flow location of the control point; and removing the control point from the plurality of control points when the calculated distance is greater than a threshold value.
11 . The non-transitory computer-readable medium of claim 9 , wherein the instructions to generate the warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence further comprise:
determining first control point location data for the first video frame and second control point location data for a target video frame of the video sequence from the generated control point location data; generating a warping function using the first control point location data and the second control point location data; and warping the new texture from the first video frame to the target video frame using the generated warping function.
12 . The non-transitory computer-readable medium of claim 11 , wherein the instructions to generate the blended video frames by blending the new texture in the warped video frames further comprise:
for each video frame of the additional video frames of the video sequence:
generating a preliminary blend by blending the new texture with the region in the video frame,
computing high-frequency residuals lost during the generating of the preliminary blending by blending a transparent image with the video frame, and
computing a final blend of the new texture with the region in the video frame by merging the preliminary blend and the computed high-frequency residuals.
13 . The non-transitory computer-readable medium of claim 9 , wherein the instructions to generate the transformation function representing the deformation of the plurality of control points using the determined motions of each of the plurality of control points further comprise:
generating a first transformation function representing rigid deformation and a second transformation function representing non-rigid deformation; and generating the transformation function by applying a weighting to the first transformation function and the second transformation function.
14 . The non-transitory computer-readable medium of claim 8 , wherein the instructions to generate the blended video frames by blending the new texture in the warped video further comprise:
receiving a second input including a prompt, the prompt indicating a requested texture for the selected region of the first video frame.
15 . A system comprising:
a memory component; and a processing device coupled to the memory component, the processing device to perform operations comprising:
obtaining a new texture for a selected region of a first video frame of a video sequence;
generating a mesh for the selected region of the first video frame of the video sequence, wherein the mesh includes a plurality of control points arranged within the mesh;
determining control point location data for each of the plurality of control points for additional video frames of the video sequence;
generating warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence;
generating blended video frames by blending the new texture in the warped video frames; and
providing a modified version of the video sequence using the generated blended video frames.
16 . The system of claim 15 , wherein the operations of determining the control point location data for each of the plurality of control points for the additional video frames of the video sequence further comprise:
for each pair of consecutive video frames of the video sequence:
determining motions of each of the plurality of control points from the first video frame of the video sequence to a second video frame of the video sequence from an optical flow mapping of the video sequence,
generating a transformation function representing deformation of the plurality of control points using the determined motions of each of the plurality of control points,
generating locations for each of the plurality of control points in the second video frame of the video sequence by warping the plurality of control points from the first video frame of the video sequence using the generated transformation function, and
storing the generated locations for each of the plurality of control points in the second video frame of the video sequence as the control point location data.
17 . The system of claim 16 , wherein the processing device performs further operations comprising:
determining a reverse optical flow location of a control point in the first video frame of the video sequence using a reverse optical flow mapping from the second video frame of the video sequence to the first video frame of the video sequence; calculating a distance between an original location of the control point and a reverse optical flow location of the control point; and removing the control point from the plurality of control points when the calculated distance is greater than a threshold value.
18 . The system of claim 16 , wherein the operations of generating the warped video frames by applying the new texture to the additional video frames of the video sequence using the control point location data for the additional video frames of the video sequence further comprise:
determining first control point location data for the first video frame and second control point location data for a target video frame of the video sequence from the generated control point location data; generating a warping function using the first control point location data and the second control point location data; and warping the new texture from the first video frame to the target video frame using the generated warping function.
19 . The system of claim 18 , wherein the operations of generating the blended video frames by blending the new texture in the warped video frames further comprise:
for each video frame of the additional video frames of the video sequence:
generating a preliminary blend by blending the new texture with the region in the video frame,
computing high-frequency residuals lost during the generating of the preliminary blending by blending a transparent image with the video frame, and
computing a final blend of the new texture with the region in the video frame by merging the preliminary blend and the computed high-frequency residuals.
20 . The system of claim 16 , wherein the operations of generating the transformation function representing the deformation of the plurality of control points using the determined motions of each of the plurality of control points comprise:
generating a first transformation function representing rigid deformation and a second transformation function representing non-rigid deformation; and generating the transformation function by applying a weighting to the first transformation function and the second transformation function.Join the waitlist — get patent alerts
Track US2025278868A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.