Automatic generation of composite images
Abstract
Certain aspects and features of this disclosure relate to automatic generation of composite images. For example, a method involves producing a representative image corresponding to a composite image based on a presentation context of input objects and segmenting the generated objects from the representative image to extract the generated objects from the representative image. The method also includes generating an inferred disposition of each of the generated objects and transforming each of the input objects to the inferred disposition of a corresponding generated object. The method can also include transmitting, storing, display, or rendering, in response to the transforming, the composite image of the input objects. Certain aspects and features also include computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
producing, using a generative image model, a representative image corresponding to a composite image based on a presentation context of a plurality of input objects, the representative image including a plurality of generated objects corresponding to the plurality of input objects; segmenting, using an image mask module, the plurality of generated objects from the representative image to extract the plurality of generated objects from the representative image; generating, using an image overlay module, an inferred disposition of each of the plurality of generated objects as extracted from the representative image; transforming, using the image overlay module, each the plurality of input objects to the inferred disposition of a corresponding generated object from the plurality of generated objects; and rendering, in response to the transforming, the composite image of the plurality of input objects.
2 . The method of claim 1 , further comprising inpainting the composite image to fill any spaces in the composite image.
3 . The method of claim 1 , wherein producing the representative image further comprises:
optimizing, using the inferred disposition, an image generation prompt; and providing the image generation prompt as optimized to the generative image model to produce the representative image.
4 . The method of claim 3 , further comprising:
generating a vector corresponding to an embedding space that represents the presentation context; sampling the embedding space around the vector; and optimizing the image generation prompt at least in part in response to the sampling of the embedding space.
5 . The method of claim 1 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
6 . The method of claim 1 , further comprising inferring, using an object inference module, a location corresponding to each of the plurality of input objects in a plurality of object images.
7 . The method of claim 6 , wherein each of the plurality of input objects is described by metadata associated with an image from among the plurality of object images, the method further comprising:
identifying each of the plurality of input objects based in part on the metadata associated with the image; and inferring the location corresponding to each of the plurality of input objects based at least in part on the metadata.
8 . A system comprising:
a memory component; a processing device coupled to the memory component to perform operations of receiving a plurality of input images including input objects and causing a composite image of the input objects to be transmitted, stored, or displayed: a generative image model configured to produce a representative image corresponding to the composite image based on a presentation context of the input objects, the representative image including a plurality of generated objects corresponding to the input objects; an image mask module configured to segment the plurality of generated objects from the representative image to extract the plurality of generated objects from the representative image; and an image overlay module configured to transform each the input objects to an inferred disposition of a corresponding generated object from the plurality of generated objects to produce the composite image of the input objects.
9 . The system of claim 8 , further comprising an inpainting module configured to fill any spaces in the composite image.
10 . The system of claim 8 , further comprising an object inference module configured to infer the input objects from the plurality of input images.
11 . The system of claim 8 , further comprising a prompt generation module configured to sample an embedding space around a vector that represents the presentation context and optimize an image generation prompt to cause the generative image model to produce the representative image.
12 . The system of claim 8 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
13 . The system of claim 8 , wherein each of the input objects is described by metadata associated with an image from the plurality of input images, and wherein the metadata configured to identify the input objects and the presentation context of the input objects.
14 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:
producing, using a generative image model, a representative image corresponding to a composite image based on a presentation context of a plurality of input objects; a step for transforming each the plurality of input objects to an inferred disposition of a generated object in the representative image to produce the composite image of the plurality of input objects; and rendering the composite image of the plurality of input objects.
15 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations comprising inpainting the composite image to fill any spaces in the composite image.
16 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations comprising:
optimizing, using the inferred disposition, an image generation prompt; and providing the image generation prompt as optimized to the generative image model to produce the representative image.
17 . The non-transitory computer-readable medium of claim 16 , wherein the instructions further cause the processing device to perform operations comprising:
generating a vector corresponding to an embedding space that represents the presentation context; sampling the embedding space around the vector; and optimizing the image generation prompt at least in part in response to the sampling of the embedding space.
18 . The non-transitory computer-readable medium of claim 14 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
19 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations comprising inferring a location corresponding to each of the plurality of input objects in a plurality of object images.
20 . The non-transitory computer-readable medium of claim 19 , wherein each of the plurality of input objects is described by metadata associated with an image from among a plurality of input images, and wherein the instructions further cause the processing device to perform operations comprising:
identifying each of the plurality of input objects based in part on the metadata associated with the image; and inferring the location corresponding to each of the plurality of input objects based at least in part on the metadata.Join the waitlist — get patent alerts
Track US2024386633A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.