US2026050826A1PendingUtilityA1

Dynamically managing prompts and model parameters for high-throughput text-to-image inference serving

Assignee: ADOBE INCPriority: Aug 19, 2024Filed: Aug 19, 2024Published: Feb 19, 2026
Est. expiryAug 19, 2044(~18.1 yrs left)· nominal 20-yr term from priority
G06N 3/045G06T 11/00G06N 20/00
65
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure describes one or more implementations of systems that utilizes prompt-aware, accuracy-scaling inference serving to serve prompts into generative models. For instance, the disclosed systems utilize varying approximation levels in generative models to speed up model output inferences. For example, the disclosed systems determine a set of approximation parameters for a set of generative models based on a predicted input prompt load. The disclosed systems generate an input prompt distribution mapping utilizing a historical prompt affinity mapping to the set of generative models and a prompt load distribution for the set of generative models. The disclosed systems select, for an input prompt, a generative model corresponding to a particular approximation parameter based on the input prompt distribution mapping and an approximation parameter assignment for the input prompt. The disclosed systems generate an inference output for the input prompt by utilizing the input prompt with the generative model.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method comprising:
 determining a set of approximation parameters for a set of generative models based on a predicted input prompt load;   generating an input prompt distribution mapping utilizing a historical prompt affinity mapping to the set of generative models and a prompt load distribution for the set of generative models;   selecting, for an input prompt, a generative model corresponding to a particular approximation parameter based on the input prompt distribution mapping and an approximation parameter assignment for the input prompt; and   generating an inference output for the input prompt by utilizing the input prompt with the generative model.   
     
     
         2 . The computer-implemented method of  claim 1 , further comprising determining an approximation parameter for the generative model by configuring a number of skipped denoising iterations for the generative model based on the predicted input prompt load. 
     
     
         3 . The computer-implemented method of  claim 1 , further comprising determining the historical prompt affinity mapping to the set of generative models by determining, for a historical prompt, a target approximation parameter based on image quality scores corresponding to output images generated by the set of generative models for the historical prompt. 
     
     
         4 . The computer-implemented method of  claim 1 , further comprising determining a load distribution for the set of generative models corresponding to the set of approximation parameters by determining a fraction of input prompts to process at each generative model from the set of generative models to satisfy a throughput target. 
     
     
         5 . The computer-implemented method of  claim 1 , further comprising generating the input prompt distribution mapping by determining prompt shift probabilities that represent redirection probabilities for historical prompts to particular generative models based on target approximation parameters of the historical prompts from the historical prompt affinity mapping fitting the prompt load distribution. 
     
     
         6 . The computer-implemented method of  claim 1 , further comprising determining an approximation parameter assignment for the input prompt by:
 identifying a historical prompt for the input prompt based on a similarity score between the historical prompt and the input prompt; and   mapping a target approximation parameter corresponding to the historical prompt to the input prompt.   
     
     
         7 . The computer-implemented method of  claim 1 , further comprising:
 selecting, for an additional input prompt, the generative model corresponding to the particular approximation parameter based on the input prompt distribution mapping and an additional approximation parameter assignment for the additional input prompt; and   upon determining that the input prompt and the additional input prompt satisfies an input prompt load threshold, generating a batch inference output by utilizing the input prompt and the additional input prompt as a batch of input prompts with the generative model.   
     
     
         8 . The computer-implemented method of  claim 1 , wherein:
 the set of generative models comprise a set of text-to-image diffusion models;   the input prompt comprises a text prompt requesting an image based on a description in the text prompt; and   the inference output comprises an output image.   
     
     
         9 . The computer-implemented method of  claim 8 , further comprising:
 identifying an updated predicted input prompt load;   determining an updated set of approximation parameters for the set of generative models;   generating an updated input prompt distribution mapping based on the updated predicted input prompt load; and   selecting, for an additional input prompt, an additional generative model corresponding to an additional particular approximation parameter based on the updated input prompt distribution mapping.   
     
     
         10 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
 identifying a set of generative models corresponding to different approximation parameters;   generating an input prompt distribution mapping by:
 identifying a prompt load distribution for the set of generative models; 
 determining a historical prompt affinity mapping to the set of generative models utilizing affinities between historical prompts and generative models from the set of generative models; and 
 determining a prompt shift probability utilizing the historical prompt affinity mapping and the prompt load distribution; 
   identifying an input prompt requesting content generation through generative models, wherein the input prompt corresponds to an approximation parameter assignment; and   selecting, from the set of generative models, a generative model for the input prompt by utilizing the input prompt distribution mapping and the approximation parameter assignment for the input prompt.   
     
     
         11 . The non-transitory computer-readable medium of  claim 10 , wherein the operations further comprise determining the different approximation parameters by modifying a set of approximation parameters corresponding to the set of generative models based on a predicted input prompt load, wherein an approximation parameter comprises a number of skipped denoising iterations. 
     
     
         12 . The non-transitory computer-readable medium of  claim 10 , wherein the operations further comprise identifying the prompt load distribution by determining a fraction of input prompts to process at each generative model from the set of generative models. 
     
     
         13 . The non-transitory computer-readable medium of  claim 10 , wherein the operations further comprise determining the approximation parameter assignment for the input prompt by:
 identifying a historical prompt for the input prompt based on a similarity score between the historical prompt and the input prompt; and   mapping a target approximation parameter corresponding to the historical prompt to the input prompt.   
     
     
         14 . The non-transitory computer-readable medium of  claim 13 , wherein the operations further comprise selecting the generative model for the input prompt by:
 determining a redirected approximation parameter for the input prompt based on an availability of the approximation parameter assignment in the input prompt distribution mapping; and   selecting the generative model, from the set of generative models, corresponding to the redirected approximation parameter.   
     
     
         15 . The non-transitory computer-readable medium of  claim 10 , wherein the operations further comprise determining the prompt shift probability by determining redirection probabilities for the historical prompts to particular generative models based on target approximation parameters of the historical prompts from the historical prompt affinity mapping fitting the prompt load distribution. 
     
     
         16 . A system comprising:
 a memory component comprising a set of generative models corresponding to different approximation parameters; and   a processing device coupled to the memory component, the processing device to perform operations comprising:
 generating an input prompt distribution mapping utilizing a historical prompt affinity mapping to the set of generative models and a prompt load distribution for the set of generative models; and 
 utilizing an input prompt with a generative model, from the set of generative models, by:
 determining an approximation parameter assignment for the input prompt based on a similarity between the input prompt and a historical input prompt comprising a target approximation parameter; and 
 selecting the generative model corresponding to an additional target approximation parameter for the input prompt based on the input prompt distribution mapping and the target approximation parameter. 
 
   
     
     
         17 . The system of  claim 16 , wherein the set of generative models comprise a set of text-to-image diffusion models and wherein the operations further comprise determining the different approximation parameters by configuring a number of skipped denoising iterations for the set of text-to-image diffusion models. 
     
     
         18 . The system of  claim 16 , wherein the operations further comprise determining the input prompt distribution mapping by determining prompt shift probabilities that indicate redirection probabilities for historical prompts based on target approximation parameters of the historical prompts from the historical prompt affinity mapping fitting the prompt load distribution. 
     
     
         19 . The system of  claim 18 , wherein the operations further comprise determining the additional target approximation parameter for the input prompt based on an availability of the target approximation parameter in the input prompt distribution mapping. 
     
     
         20 . The system of  claim 19 , wherein the operations further comprise generating an inference output for the input prompt by utilizing the input prompt with the generative model at an approximation level corresponding to the additional target approximation parameter.

Join the waitlist — get patent alerts

Track US2026050826A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.