US2025045622A1PendingUtilityA1
Batch scheduling for efficient execution of multiple machine learning models
Est. expiryAug 4, 2043(~17 yrs left)· nominal 20-yr term from priority
G06N 20/00
60
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Apparatuses, systems, and techniques for efficient profiling, scheduling, and batch execution of multiple machine learning models (MLMs). Efficient batch execution includes obtaining execution metrics characterizing expected utilization of computational resources by the MLMs, and generating at least one batch queue having one or more MLM batches of MLMs with a combined expected utilization not exceeding a threshold utilization, and initiating parallel execution of the MLMs using the generated MLM batches.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving an identification of a plurality of machine learning models (MLMs) for execution on a set of computational resources; obtaining execution metrics characterizing expected utilization of the set of computational resources during execution of individual MLMs of the plurality of MLMs; generating a first batch queue comprising one or more MLM batches, wherein at least one MLM batch comprises one or more MLMs of the plurality of MLMs, at least one MLM batch having a combined expected utilization of the set of computational resources not exceeding a threshold utilization; and initiating parallel execution of a first MLM batch of the one or more MLM batches of the first batch queue.
2 . The method of claim 1 , wherein the execution metrics characterizing expected utilization of the set of computational resources include at least one of:
a size of input data into an MLM of the plurality of MLMs, a total memory used during execution of the MLM, a peak memory use during execution of the MLM, or a peak processing clock speed during execution of the MLM.
3 . The method claim 1 , wherein the execution metrics further include expected utilization of one or more virtual processing units supported by the set of computational resources.
4 . The method of claim 1 , wherein obtaining the execution metrics for an MLM of the plurality of MLMs comprises:
collecting the execution metrics during individual execution of the MLM.
5 . The method of claim 4 , further comprising:
storing the collected execution metrics in a memory device.
6 . The method of claim 1 , wherein obtaining the execution metrics for an MLM of the plurality of MLMs comprises estimating the execution metrics for the MLM using one or more of:
an architecture of the MLM, a size of an input into the MLM, a number of computational operations associated with the MLM, or one or more number formats used by the computational operations associated with the MLM.
7 . The method of claim 1 , wherein the combined expected utilization of the set of computational resources by the first MLM batch characterizes expected utilization of memory resources during parallel execution of one or more MLMs of the first MLM batch.
8 . The method of claim 7 , wherein the combined expected utilization of the set of computational resources by the first MLM batch characterizes expected utilization of one or more processing units during parallel execution of the one or more MLMs of the first MLM batch.
The method of claim 1 , wherein of the set of computational resources comprises at least one of a central processing unit (CPU), a data processing unit (DPU), or a graphics processing unit (GPU).
9 . The method of claim 1 , further comprising:
initiating, concurrently with the parallel execution of the first MLM batch, parallel execution of a second MLM batch of the one or more MLM batches of the first batch queue, the first MLM batch and the second MLM batch being executed on:
two or more separate graphics processing units (GPUs), or
two or more separate virtual GPUs supported by a same GPU.
10 . The method of claim 1 , further comprising:
responsive to completing execution of the first MLM batch, initiating parallel execution of a second MLM batch of the one or more MLM batches of the first batch queue.
11 . The method of claim 1 , further comprising:
subsequent to initiating execution of the first MLM batch, generating at least a second batch queue, wherein the second batch queue comprises at least one MLM batch that is different from at least one other MLM batch of the first batch queue.
12 . The method of claim 12 , further comprising:
determining first performance metrics associated with execution of the first batch queue; computing second performance metrics associated with prospective execution of the second batch queue; and responsive to a comparison of the first performance metrics and the second performance metrics, switching from the execution of the first batch queue to an execution of the second batch queue.
13 . The method of claim 12 , further comprising:
displaying a first efficiency report to a user, wherein the first efficiency report comprises runtime performance metrics associated with execution of the first batch queue; displaying a second efficiency report to the user, wherein the second efficiency report comprises estimated performance metrics associated with prospective execution of the second batch queue; and responsive to receiving, from the user, a selection of the second batch queue, switching from execution of the first batch queue to execution of the second batch queue.
14 . The method of claim 12 , further comprising:
storing at least one of the first batch queue or the second batch queue in a memory device.
15 . The method of claim 1 , wherein generating the first batch queue comprises:
forming, using a priority metric, a priority queue for the plurality of MLMs; and performing a plurality of MLM placement operations, wherein individual MLM placement operations comprise:
selecting a next MLM in the priority queue;
placing the selected MLM, using the threshold utilization and the execution metrics for the selected MLM, into at least one of:
an existing MLM batch of the first batch queue, or
a new MLM batch of the first batch queue.
16 . A system comprising:
a memory device; and a processor, communicatively coupled to the memory device, to:
receive an identification of a plurality of machine learning models (MLMs) for execution on a set of computational resources;
obtain execution metrics characterizing expected utilization of the set of computational resources during execution of individual MLMs of the plurality of MLMs;
generating a first batch queue comprising one or more MLM batches, wherein each MLM batch comprises one or more MLMs of the plurality of MLMs, each MLM batch having a combined expected utilization of the set of computational resources not exceeding a threshold utilization; and initiate parallel execution of a first MLM batch of the one or more MLM batches of the first batch queue.
17 . The system of claim 17 , wherein to obtain the execution metrics for an MLM of the plurality of MLMs, the processing device is to perform at least one of:
collect the execution metrics during individual execution of the MLM; or. estimate the execution metrics for the MLM using one or more of:
an architecture of the MLM,
a size of an input into the MLM,
a number of computational operations associated with the MLM, or
one or more number formats used by the computational operations associated with the MLM.
18 . The system of claim 17 , wherein the combined expected utilization of the set of computational resources by the first MLM batch characterizes at least one of:
an expected utilization of memory resources during parallel execution of one or more MLMs of the first MLM batch, or an expected utilization of one or more processing units during parallel execution of the one or more MLMs of the first MLM batch.
19 . The system of claim 17 , wherein the processing device is further to:
subsequent to initiating execution of the first MLM batch, generate at least a second batch queue, wherein the second batch queue comprises at least one MLM batch that is different from each MLM batch of the first batch queue; determine first performance metrics associated with execution of the first batch queue; compute second performance metrics associated with prospective execution of the second batch queue; and responsive to a comparison of the first performance metrics and the second performance metrics, switching from the execution of the first batch queue to an execution of the second batch queue.
20 . The system of claim 17 , wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using one or more application programming interfaces; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
21 . A processor comprising processing circuitry to perform operations comprising:
receiving an identification of a plurality of machine learning models (MLMs) for execution on a set of computational resources; obtaining execution metrics characterizing expected utilization of the set of computational resources during execution of individual MLMs of the plurality of MLMs; generating a first batch queue comprising one or more MLM batches, wherein at least one MLM batch comprises one or more MLMs of the plurality of MLMs, at least one MLM batch having a combined expected utilization of the set of computational resources not exceeding a threshold utilization; and initiating parallel execution of a first MLM batch of the one or more MLM batches of the first batch queue.Join the waitlist — get patent alerts
Track US2025045622A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.