US2025272807A1PendingUtilityA1

Mask-free composite image generation

Assignee: ADOBE INCPriority: Feb 22, 2024Filed: Feb 22, 2024Published: Aug 28, 2025
Est. expiryFeb 22, 2044(~17.6 yrs left)· nominal 20-yr term from priority
G06N 3/0455G06T 2207/20221G06T 2207/20084G06T 2207/20081G06N 3/09G06N 3/084G06N 3/098G06N 3/0475G06N 3/0464G06T 7/194G06T 7/11G06T 5/50H04N 5/272G06T 5/60G06N 3/094G06T 11/60G06T 7/149G06T 5/75G06T 5/77
62
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, apparatus, non-transitory computer readable medium, and system include obtaining a first image depicting a background scene and a second image depicting a foreground element, generating a guidance embedding based on the second image, and generating a synthetic image depicting the foreground element and the background scene based on the first image and the guidance embedding, wherein the image generation model determines a location of the foreground element within the synthetic image in light of the background scene.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 obtaining a first image depicting a background scene and a second image depicting a foreground element;   generating, using an adapter network, a guidance embedding based on the second image; and   generating, using an image generation model, a synthetic image depicting the foreground element and the background scene based on the first image and the guidance embedding, wherein the image generation model determines a location of the foreground element within the synthetic image in light of the background scene.   
     
     
         2 . The method of  claim 1 , wherein generating the guidance embedding comprises:
 encoding, using an image encoder, the second image to obtain an image embedding, wherein the guidance embedding is generated based on the image embedding.   
     
     
         3 . The method of  claim 1 , generating the synthetic image comprises:
 combining the first image with a noise map to obtain a noisy image, wherein the image generation model takes the noisy image as input.   
     
     
         4 . The method of  claim 3 , wherein:
 the noise map having a same resolution as the first image.   
     
     
         5 . The method of  claim 1 , generating the synthetic image comprises:
 performing a reverse diffusion process.   
     
     
         6 . The method of  claim 1 , further comprising:
 obtaining location information for the foreground element of the second image, wherein a location of the foreground element in the synthetic image is determined based on the location information.   
     
     
         7 . The method of  claim 6 , wherein:
 a size of the foreground element is determined by the location information.   
     
     
         8 . The method of  claim 1 , further comprising:
 obtaining an additional image depicting an additional foreground element, wherein the synthetic image is generated to depict the additional foreground element.   
     
     
         9 . The method of  claim 1 , wherein:
 the image generation model is trained to combine multiple images using training data that includes a foreground image, a background image, and a ground truth image that combines the foreground image and the background image.   
     
     
         10 . A method comprising:
 obtaining training data including a training foreground image, a training background image, and a ground truth image that combines the training foreground image and the training background image; and   training, using the training data, an image generation model to generate a synthetic image based on an input foreground image and an input background image, wherein the synthetic image includes a foreground element from the input foreground image and a background region from the input background image.   
     
     
         11 . The method of  claim 10 , wherein obtaining the training data comprises:
 identifying a region of the ground truth image depicting the foreground element; and   performing object segmentation on the ground truth image to obtain a segmentation mask, wherein the training foreground image and the training background image are based on the segmentation mask.   
     
     
         12 . The method of  claim 11 , further comprising:
 generating an inpainting mask based on the segmentation mask, wherein the inpainting mask is different than the segmentation mask and the training background image is based on the inpainting mask.   
     
     
         13 . The method of  claim 12 , further comprising:
 performing inpainting based on the inpainting mask to obtain an inpainted image, wherein the training background image is based on the inpainted image.   
     
     
         14 . The method of  claim 13 , further comprising:
 refining the inpainted image to obtain the training background image.   
     
     
         15 . The method of  claim 12 , wherein:
 the inpainting mask includes a shadow region or a reflection region that is absent from the segmentation mask.   
     
     
         16 . The method of  claim 10 , wherein the training the image generation model comprises:
 generating, using an adapter network, a guidance embedding based on the training foreground image;   generating, using the image generation model, a predicted image based on the training background image and the guidance embedding;   computing a loss function based on the predicted image and the ground truth image; and   updating parameters of the adapter network based on the loss function.   
     
     
         17 . An apparatus comprising:
 at least one processor;   at least one memory storing instructions executable by the at least one processor;   an image encoder comprising parameters stored in the at least one memory and trained to encode a foreground image to obtain an image embedding;   an adapter network comprising parameters stored in the at least one memory and trained to generate a guidance embedding based on the image embedding; and   an image generation model comprising parameters stored in the at least one memory and trained to generate a synthetic image based on a background image and the guidance embedding, wherein the synthetic image depicts a foreground element and a first portion of the background scene, and wherein the foreground element is located at a position of the synthetic image that is determined by the image generation model.   
     
     
         18 . The apparatus of  claim 17 , further comprising:
 an object detection component comprising parameters stored in the at least one memory and trained to identifying the foreground element from an image.   
     
     
         19 . The apparatus of  claim 17 , further comprising:
 a segmentation component comprising parameters stored in the at least one memory and trained to perform object segmentation on an image to generate a segmentation mask.   
     
     
         20 . The apparatus of  claim 19 , further comprising:
 an inpainting component comprising parameters stored in the at least one memory and trained to generate the background image based on the segmentation mask.

Join the waitlist — get patent alerts

Track US2025272807A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.