US2025054211A1PendingUtilityA1

Techniques for using machine learning models to generate scenes based upon image tiles

Assignee: AUTODESK INCPriority: Aug 11, 2023Filed: Apr 29, 2024Published: Feb 13, 2025
Est. expiryAug 11, 2043(~17 yrs left)· nominal 20-yr term from priority
G06T 11/23G06N 5/04G06N 20/00G06T 11/00G06T 7/11G06T 11/203
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In various embodiments, examples of the disclosure provide systems and methods for generating a scene image. A plurality of input image tiles based upon at least one user input are obtained. Spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles is detected, and then a scene composition of the scene determined based upon the spatial positioning of each input image tile included in the plurality of image tiles. A scene prompt associated with the scene is obtained, and a machine learning model blends the plurality of image tiles based upon the scene prompt to generate the scene.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for generating a scene, the method comprising:
 obtaining a plurality of input image tiles based upon at least one user input;   detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles;   determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles;   obtaining a scene prompt associated with the scene; and   causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein the scene prompt comprises a textual prompt indicating how the plurality of image tiles should be blended to generate the scene. 
     
     
         3 . The computer-implemented method of  claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
 obtaining an image tile textual prompt that describes a desired characteristic of a corresponding image tile included in the plurality of image tiles;   providing the image tile textual prompt to the machine learning model; and   executing the machine learning model to generate the corresponding image tile based upon the textual prompt.   
     
     
         4 . The computer-implemented method of  claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
 obtaining within a graphical user interface a sketch of a portion of the scene;   providing the sketch to the machine learning model; and   executing the machine learning model to generate a corresponding image tile based upon the sketch.   
     
     
         5 . The computer-implemented method of  claim 1 , further comprising:
 obtaining within a graphical user interface a sketch of a portion of the scene;   obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile; and   executing the machine learning model to generate the respective image tile based upon the sketch and the image tile textual prompt.   
     
     
         6 . The computer-implemented method of  claim 1 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
 obtaining within a graphical user interface a designation of a region of a corresponding image tile;   obtaining a region textual prompt that corresponds to the region, wherein the region textual prompt specifies a desired characteristic of the region within the respective image tile; and   executing the machine learning model to generate the corresponding image tile based upon the region and the region textual prompt.   
     
     
         7 . The computer-implemented method of  claim 1 , wherein obtaining the plurality of input image tiles based upon at least one user input comprises:
 obtaining a selection of a previously generated image tile included in the plurality of previously generated image tiles;   generating a new image tile based upon a user input;   obtaining a textual prompt specifying a desired characteristic for a combined image tile; and   executing the machine learning model to generate the combined image tile based upon the desired characteristic, the previously generated image tile, and the new image tile.   
     
     
         8 . The computer-implemented method of  claim 1 , further comprising:
 obtaining, within a graphical user interface, the spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles, wherein the spatial positioning comprises whitespace between at least a first input image tile included in the plurality of image tiles and a second input image tile included in the plurality of input image tiles.   
     
     
         9 . The computer-implemented method of  claim 8 , further comprising automatically adjusting an amount of whitespace between each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles. 
     
     
         10 . The computer-implemented method of  claim 8 , wherein executing the first machine learning model to blend the plurality of image tiles to form the scene based upon the scene prompt comprises filling the whitespace between the at least one of the plurality of image tiles based upon the scene prompt. 
     
     
         11 . One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to generate a scene by performing the steps of:
 obtaining a plurality of input image tiles based upon at least one user input;   detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles;   determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles;   obtaining a scene prompt associated with the scene; and   causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.   
     
     
         12 . The one or more non-transitory computer readable media of  claim 11 , wherein the scene prompt comprises a textual prompt indicating how the plurality of image tiles should be blended to generate the scene. 
     
     
         13 . The one or more non-transitory computer readable media of  claim 11 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
 obtaining an image tile textual prompt that describes a desired characteristic of a corresponding image tile included in the plurality of image tiles;   providing the image tile textual prompt to the machine learning model; and   executing the machine learning model to generate the corresponding image tile based upon the textual prompt.   
     
     
         14 . The one or more non-transitory computer readable media of  claim 11 , wherein the method comprises:
 obtaining within a graphical user interface a sketch of a portion of the scene;   providing the sketch to the machine learning model; and   executing the machine learning model to generate a corresponding image tile based upon the sketch.   
     
     
         15 . The one or more non-transitory computer readable media of  claim 14 , further comprising:
 obtaining within a graphical user interface a sketch of a portion of the scene;   obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile; and   executing the machine learning model to generate the respective image tile based upon the sketch and the image tile textual prompt.   
     
     
         16 . The one or more non-transitory computer readable media of  claim 11 , wherein obtaining the plurality of input image tiles based upon the at least one user input comprises:
 obtaining within a graphical user interface a designation of a region of a corresponding image tile;   obtaining a region textual prompt that corresponds to the region, wherein the region textual prompt specifies a desired characteristic of the region within the respective image tile; and   executing the machine learning model to generate the corresponding image tile based upon the region and the region textual prompt.   
     
     
         17 . The one or more non-transitory computer readable media of  claim 11  wherein obtaining the plurality of input image tiles based upon at least one user input comprises:
 obtaining a selection of a previously generated image tile included in the plurality of previously generated image tiles; 
 generating a new image tile based upon a user input; 
 obtaining a textual prompt specifying a desired characteristic for a combined image tile; and 
 executing the machine learning model to generate the combined image tile based upon the desired characteristic, the previously generated image tile, and the new image tile. 
 
     
     
         18 . The one or more non-transitory computer readable media of  claim 11 , wherein the method comprises:
 obtaining, within a graphical user interface, the spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles, wherein the spatial positioning comprises whitespace between at least a first input image tile included in the plurality of image tiles and a second input image tile included in the plurality of input image tiles.   
     
     
         19 . The one or more non-transitory computer readable media of  claim 18 , wherein causing the machine learning model to blend the plurality of image tiles to form the scene based upon the scene prompt comprises filling the whitespace between the at least one of the plurality of image tiles based upon the scene prompt. 
     
     
         20 . A system comprising:
 one or more memories storing instructions; and
 one or more processors coupled to the one or more memories that, when executed, perform the steps of:
 obtaining a plurality of input image tiles based upon at least one user input; 
 detecting a spatial positioning of each input image tile included in the plurality of input image tiles relative to at least one other input image tile included in the plurality of input image tiles; 
 determining a scene composition of the scene based upon the spatial positioning of each input image tile included in the plurality of image tiles; 
 obtaining a scene prompt associated with the scene; and 
 causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene.

Join the waitlist — get patent alerts

Track US2025054211A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.