Techniques for using machine learning models to generate scenes based upon image tiles
Abstract
In various embodiments, examples of the disclosure provide systems and methods for generating a scene image. A plurality of input image tiles based upon at least one user input are obtained. Spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles is detected, and then a scene composition of the scene determined based upon the spatial positioning of each input image tile included in the plurality of image tiles. A scene prompt associated with the scene is obtained, and a machine learning model blends the plurality of image tiles based upon the scene prompt to generate the scene.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for generating a scene, the method comprising:
obtaining a plurality of input image tiles based upon at least one user input; detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles; determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles; obtaining a scene prompt associated with the scene; and causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.
2 . The computer-implemented method of claim 1 , wherein the scene prompt comprises a textual prompt indicating how the plurality of image tiles should be blended to generate the scene.
3 . The computer-implemented method of claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
obtaining an image tile textual prompt that describes a desired characteristic of a corresponding image tile included in the plurality of image tiles; providing the image tile textual prompt to the machine learning model; and executing the machine learning model to generate the corresponding image tile based upon the textual prompt.
4 . The computer-implemented method of claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
obtaining within a graphical user interface a sketch of a portion of the scene; providing the sketch to the machine learning model; and executing the machine learning model to generate a corresponding image tile based upon the sketch.
5 . The computer-implemented method of claim 1 , further comprising:
obtaining within a graphical user interface a sketch of a portion of the scene; obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile; and executing the machine learning model to generate the respective image tile based upon the sketch and the image tile textual prompt.
6 . The computer-implemented method of claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
obtaining within a graphical user interface a designation of a region of a corresponding image tile; obtaining a region textual prompt that corresponds to the region, wherein the region textual prompt specifies a desired characteristic of the region within the respective image tile; and executing the machine learning model to generate the corresponding image tile based upon the region and the region textual prompt.
7 . The computer-implemented method of claim 1 , wherein obtaining the plurality of input image tiles based upon at least one user input comprises:
obtaining a selection of a previously generated image tile included in the plurality of previously generated image tiles; generating a new image tile based upon a user input; obtaining a textual prompt specifying a desired characteristic for a combined image tile; and executing the machine learning model to generate the combined image tile based upon the desired characteristic, the previously generated image tile, and the new image tile.
8 . The computer-implemented method of claim 1 , further comprising:
obtaining, within a graphical user interface, the spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles, wherein the spatial positioning comprises whitespace between at least a first input image tile included in the plurality of image tiles and a second input image tile included in the plurality of input image tiles.
9 . The computer-implemented method of claim 8 , further comprising automatically adjusting an amount of whitespace between each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles.
10 . The computer-implemented method of claim 8 , wherein executing the first machine learning model to blend the plurality of image tiles to form the scene based upon the scene prompt comprises filling the whitespace between the at least one of the plurality of image tiles based upon the scene prompt.
11 . One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to generate a scene by performing the steps of:
obtaining a plurality of input image tiles based upon at least one user input; detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles; determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles; obtaining a scene prompt associated with the scene; and causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.
12 . The one or more non-transitory computer readable media of claim 11 , wherein the scene prompt comprises a textual prompt indicating how the plurality of image tiles should be blended to generate the scene.
13 . The one or more non-transitory computer readable media of claim 11 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
obtaining an image tile textual prompt that describes a desired characteristic of a corresponding image tile included in the plurality of image tiles; providing the image tile textual prompt to the machine learning model; and executing the machine learning model to generate the corresponding image tile based upon the textual prompt.
14 . The one or more non-transitory computer readable media of claim 11 , wherein the method comprises:
obtaining within a graphical user interface a sketch of a portion of the scene; providing the sketch to the machine learning model; and executing the machine learning model to generate a corresponding image tile based upon the sketch.
15 . The one or more non-transitory computer readable media of claim 14 , further comprising:
obtaining within a graphical user interface a sketch of a portion of the scene; obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile; and executing the machine learning model to generate the respective image tile based upon the sketch and the image tile textual prompt.
16 . The one or more non-transitory computer readable media of claim 11 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
obtaining within a graphical user interface a designation of a region of a corresponding image tile; obtaining a region textual prompt that corresponds to the region, wherein the region textual prompt specifies a desired characteristic of the region within the respective image tile; and executing the machine learning model to generate the corresponding image tile based upon the region and the region textual prompt.
17 . The one or more non-transitory computer readable media of claim 11 wherein obtaining the plurality of input image tiles based upon at least one user input comprises:
obtaining a selection of a previously generated image tile included in the plurality of previously generated image tiles;
generating a new image tile based upon a user input;
obtaining a textual prompt specifying a desired characteristic for a combined image tile; and
executing the machine learning model to generate the combined image tile based upon the desired characteristic, the previously generated image tile, and the new image tile.
18 . The one or more non-transitory computer readable media of claim 11 , wherein the method comprises:
obtaining, within a graphical user interface, the spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles, wherein the spatial positioning comprises whitespace between at least a first input image tile included in the plurality of image tiles and a second input image tile included in the plurality of input image tiles.
19 . The one or more non-transitory computer readable media of claim 18 , wherein causing the machine learning model to blend the plurality of image tiles to form the scene based upon the scene prompt comprises filling the whitespace between the at least one of the plurality of image tiles based upon the scene prompt.
20 . A system comprising:
one or more memories storing instructions; and
one or more processors coupled to the one or more memories that, when executed, perform the steps of:
obtaining a plurality of input image tiles based upon at least one user input;
detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles;
determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles;
obtaining a scene prompt associated with the scene; and
causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.Join the waitlist — get patent alerts
Track US2025054211A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.