US12141622B1ActiveUtility

Dynamic distribution of a workload processing pipeline on a computing infrastructure

Assignee: ENTEFY INCPriority: Dec 29, 2017Filed: Apr 18, 2023Granted: Nov 12, 2024
Est. expiryDec 29, 2037(~11.4 yrs left)· nominal 20-yr term from priority
G06F 9/5005G06F 2209/501G06N 3/02G06T 1/20H04L 47/82G06F 9/48H04L 47/83Y02D10/00G06F 9/5027G06F 9/5094G06F 9/5044G06F 9/5077G06N 20/00
75
PatentIndex Score
0
Cited by
26
References
20
Claims

Abstract

Disclosed are systems, methods, and computer readable media for automatically assessing and allocating virtualized resources (such as central processing unit (CPU) and graphics processing unit (GPU) resources). In some embodiments, this method involves a computing infrastructure receiving a request to perform a workload, determining one or more workflows for performing the workload, selecting a virtualized resource, from a plurality of virtualized resources, wherein the virtualized resource is associated with a hardware configuration, and wherein selecting the virtualized resources is based on a suitability score determined based on benchmark scores of the one or more workflows on the hardware configuration, scheduling performance of at least part of the workload on the selected virtualized resource, and outputting results of the at least part of the workload.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
       1. A system for allocating virtualized resources, comprising:
 one or more non-transitory memory devices; and 
 one or more hardware processors configured to execute instructions from the one or more non-transitory memory devices to cause the system to perform operations comprising:
 receiving a workflow comprising a set of artificial intelligence (AI) models executable by a plurality of virtualized resources; 
 determining, based on the set of AI models, a hardware requirement to run the workflow and a feature index indicating a memory requirement to run the workflow; 
 determining hardware configuration information for host computing devices available for the plurality of virtualized resources; 
 determining a plurality of supported hardware configurations for the plurality of virtualized resources that execute the workflow based on the hardware requirement, the feature index, and the hardware configuration information; 
 running, for each of the plurality of supported hardware configurations, the set of AI models of the workflow using a corresponding one or more of the plurality of virtualized resources; 
 generating a performance benchmark for each of the plurality of supported hardware configurations based on the running; 
 determining, based on the performance benchmarks and one or more optimization parameters, a suitability score of the workflow for each of the plurality of supported hardware configurations; 
 configuring the workflow to be run on a set of the plurality of virtualized resources based on the suitability score for each of the plurality of supported hardware configurations; 
 deploying the workflow in a workflow index usable to execute the set of AI models when a corresponding request for an execution of the workflow from the workflow index is received; and 
 in response to receiving the corresponding request for the execution of the workflow from the workflow index, executing one or more application programming interface (API) calls to the set of the plurality of virtualized resources. 
 
 
     
     
       2. The system of  claim 1 , wherein the operations further comprise:
 updating, dynamically during an execution of one or more workloads using the workflow, one or more of the plurality of supported hardware configurations. 
 
     
     
       3. The system of  claim 2 , wherein, prior to the updating, the operations further comprise:
 receiving the one or more workloads from one or more user devices, wherein the one or more workloads designate ones of the plurality of virtualized resources for executing the one or more workloads based on the suitability score; and 
 executing the one or more workloads using the workflow and a plurality of additional workflows. 
 
     
     
       4. The system of  claim 1 , wherein the operations further comprise:
 updating, based on the suitability score of the workflow for each of the plurality of supported hardware configurations, a workflow index comprising available workflows that are capable of running the AI models on the plurality of virtualized resources, wherein the workflow index comprises information associated with workflow processing speeds and workflow memory size information for the available workflows when executing the AI models on the plurality of virtualized resources. 
 
     
     
       5. The system of  claim 1 , wherein the workflow comprises a plurality of sub-workflows each using one of the plurality of virtualized resources for executing a corresponding one of the AI models for parallel data processing of different regions of pixels in an image. 
     
     
       6. The system of  claim 1 , wherein the workflow applies logical operations using the set of the AI models when executed by the plurality of virtualized resources, and wherein the workload is further associated with identifying one or more objects in an image. 
     
     
       7. The system of  claim 1 , wherein the updating is based on a workflow queue metric comprising at least one of a number of additional workflows in a queue associated with executing the workflow and the additional workflows, a time in the queue for the workflow, or a workflow activity executed by a hardware resource associated with each of the plurality of supported hardware configurations. 
     
     
       8. The system of  claim 1 , wherein the suitability score for each of the supported hardware configurations is based on at least one of a percentage of resources used by a corresponding one of plurality of supported hardware configurations, an average run time of the workflow using the corresponding one of plurality of supported hardware configurations, or a power used for the corresponding one of plurality of supported hardware configurations, and wherein the suitability score further comprise a compatibility metric between each of the plurality of virtualized resources and the set of AI models. 
     
     
       9. A method for allocating computing resources, comprising:
 receiving a workflow comprising a set of the artificial intelligence (AI) models executable by a plurality of virtualized resources; 
 determining, based on the set of AI models, a hardware requirement to run the workflow and a feature index indicating a memory requirement to run the workflow; 
 determining hardware configuration information for host computing devices available for the plurality of virtualized resources; 
 determining a plurality of supported hardware configurations for the plurality of virtualized resources that execute the workflow based on the hardware requirement, the feature index, and the hardware configuration information; 
 running, for each of the plurality of supported hardware configurations, the set of AI models of the workflow using a corresponding one or more of the plurality of virtualized resources; 
 generating a performance benchmark for each of the plurality of supported hardware configurations based on the running; 
 determining, based on the performance benchmarks and one or more optimization parameters, a suitability score of the workflow for each of the plurality of supported hardware configurations; 
 configuring the workflow to be run on a set of the plurality of virtualized resources based on the suitability score for each of the plurality of supported hardware configurations; 
 deploying the workflow in a workflow index usable to execute the set of AI models when a corresponding request for an execution of the workflow from the workflow index is received; and 
 in response to receiving the corresponding request for the execution of the workflow from the workflow index, executing one or more application programming interface (API) calls to the set of the plurality of virtualized resources. 
 
     
     
       10. The method of  claim 9 , further comprising:
 updating, dynamically during an execution of one or more workloads using the workflow, one or more of the plurality of supported hardware configurations. 
 
     
     
       11. The method of  claim 10 , wherein, prior to the updating, the method further comprises:
 receiving the one or more workloads from one or more user devices, wherein the one or more workloads designate ones of the plurality of virtualized resources for executing the one or more workloads based on the suitability score; and 
 executing the one or more workloads using the workflow and a plurality of additional workflows. 
 
     
     
       12. The method of  claim 9 , further comprising;
 updating, based on the suitability score of the workflow for each of the plurality of supported hardware configurations, a workflow index comprising available workflows that are capable of running the AI models on the plurality of virtualized resources, wherein the workflow index comprises information associated with workflow processing speeds and workflow memory size information for the available workflows when executing the AI models on the plurality of virtualized resources. 
 
     
     
       13. The method of  claim 9 , wherein the workflow comprises a plurality of sub-workflows each using one of the plurality of virtualized resources for executing a corresponding one of the AI models for parallel data processing of different regions of pixels in an image. 
     
     
       14. The method of  claim 9 , wherein the workflow applies logical operations using the set of the AI models when executed by the plurality of virtualized resources, and wherein the workload is further associated with identifying one or more objects in an image. 
     
     
       15. The method of  claim 9 , wherein the updating is based on a workflow queue metric comprising at least one of a number of additional workflows in a queue associated with executing the workflow and the additional workflows, a time in the queue for the workflow, or a workflow activity executed by a hardware resource associated with each of the plurality of supported hardware configurations. 
     
     
       16. The method of  claim 9 , wherein the suitability score for each of the supported hardware configurations is based on at least one of a percentage of resources used by a corresponding one of plurality of supported hardware configurations, an average run time of the workflow using the corresponding one of plurality of supported hardware configurations, or a power used for the corresponding one of plurality of supported hardware configurations, and wherein the suitability score further comprise a compatibility metric between each of the plurality of virtualized resources and the set of AI models. 
     
     
       17. A non-transitory machine readable-medium, on which are stored instructions for allocating hardware resources, comprising instructions that when executed cause a machine to perform operations comprising:
 receiving a workflow comprising a set of the artificial intelligence (AI) models executable by a plurality of virtualized resources; 
 determining, based on the set of AI models, a hardware requirement to run the workflow and a feature index indicating a memory requirement to run the workflow; 
 determining hardware configuration information for host computing devices available for the plurality of virtualized resources; 
 determining a plurality of supported hardware configurations for the plurality of virtualized resources that execute the workflow based on the hardware requirement, the feature index, and the hardware configuration information; 
 running, for each of the plurality of supported hardware configurations, the set of AI models of the workflow using a corresponding one or more of the plurality of virtualized resources; 
 generating a performance benchmark for each of the plurality of supported hardware configurations based on the running; 
 determining, based on the performance benchmarks and one or more optimization parameters, a suitability score of the workflow for each of the plurality of supported hardware configurations; 
 configuring the workflow to be run on a set of the plurality of virtualized resources based on the suitability score for each of the plurality of supported hardware configurations; 
 deploying the workflow in a workflow index usable to execute the set of AI models when a corresponding request for an execution of the workflow from the workflow index is received; and 
 in response to receiving the corresponding request for the execution of the workflow from the workflow index, executing one or more application programming interface (API) calls to the set of the plurality of virtualized resources. 
 
     
     
       18. The non-transitory machine readable-medium of  claim 17 , wherein the operations further comprise:
 updating, dynamically during an execution of one or more workloads using the workflow, one or more of the plurality of supported hardware configurations. 
 
     
     
       19. The non-transitory machine readable-medium of  claim 18 , wherein, prior to the updating, the operations further comprise:
 receiving the one or more workloads from one or more user devices, wherein the one or more workloads designate ones of the plurality of virtualized resources for executing the one or more workloads based on the suitability score; and 
 executing the one or more workloads using the workflow and a plurality of additional workflows. 
 
     
     
       20. The non-transitory machine readable-medium of  claim 17 , wherein the operations further comprise:
 updating, based on the suitability score of the workflow for each of the plurality of supported hardware configurations, a workflow index comprising available workflows that are capable of running the AI models on the plurality of virtualized resources, wherein the workflow index comprises information associated with workflow processing speeds and workflow memory size information for the available workflows when executing the AI models on the plurality of virtualized resources.

Join the waitlist — get patent alerts

Track US12141622B1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.