Workload placement for virtual gpu enabled systems
Abstract
Disclosed are aspects of workload selection and placement in systems that include graphics processing units (GPUs) that are virtual GPU (vGPU) enabled. In some aspects, workloads are assigned to virtual graphics processing unit (vGPU)-enabled graphics processing units (GPUs). A number of vGPU placement neural networks are trained to maximize a composite efficiency metric based on workload data and GPU data for the plurality of vGPU placement models. A combined neural network selector is generated using the vGPU placement neural networks, and utilized to assign a workload to a vGPU-enabled GPU.
Claims
exact text as granted — not AI-modifiedTherefore, the following is claimed:
1 . A system comprising:
at least one computing device comprising at least one processor and at least one data store; machine readable instructions stored in the at least one data store, wherein the instructions, when executed by the at least one processor, cause the at least one computing device to at least:
identify workload data for a plurality of different datacenter configurations comprising: a plurality of different graphics processing unit (GPU) counts, a plurality of different virtual GPU (vGPU) scheduling policies, and a plurality of different workload arrival rates, wherein the workload data is generated by monitoring workloads executed or simulated using at least one computing environment configured according to the plurality of different datacenter configurations;
train a plurality of vGPU placement neural networks to maximize a composite efficiency metric based at least in part on the workload data identified for the plurality of different datacenter configurations, a respective one of the vGPU placement neural networks comprising at least two sets of layers;
generate a combined neural network selector based on the plurality of vGPU placement neural networks; and
utilize the combined neural network selector to select a particular workload of at least one candidate workload, and execute the particular workload in a particular computing environment using a particular vGPU-enabled GPU.
2 . The system of claim 1 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a workload type of the particular workload.
3 . The system of claim 1 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a datacenter configuration of the particular computing environment.
4 . The system of claim 1 , wherein monitoring the workloads comprises executing a plurality of workloads comprising a plurality of different sets of a plurality of workload parameters, wherein the plurality of workloads are executed using the at least one computing environment configured according to the plurality of different datacenter configurations.
5 . The system of claim 4 , wherein a particular parameter of the plurality of workload parameters comprises: a measure of the composite efficiency metric calculated for a workload scaled by a geometric mean of the composite efficiency metric calculated for a respective one of a plurality of workloads currently in an arrival queue awaiting selection for placement.
6 . The system of claim 4 , wherein a particular parameter of the plurality of workload parameters comprises: a workload type that indicates a purpose or activity performed.
7 . The system of claim 4 , wherein a particular parameter of the plurality of workload parameters comprises at least one of: a minimum graphics memory, and a graphics memory requested as a fraction of a total GPU memory in the at least one computing environment.
8 . A method performed by at least one computing device executing machine-readable instructions, the method comprising:
identifying workload data for a plurality of different datacenter configurations comprising at least one of: a plurality of different graphics processing unit (GPU) counts, a plurality of different virtual GPU (vGPU) scheduling policies, a plurality of different workload arrival rates, or any combination thereof, wherein the workload data is generated by monitoring workloads executed or simulated using at least one computing environment configured according to the plurality of different datacenter configurations; training a plurality of vGPU placement neural networks to maximize a composite efficiency metric based at least in part on the workload data identified for the plurality of different datacenter configurations, a respective one of the vGPU placement neural networks comprising at least two sets of layers; generating a combined neural network selector based on the plurality of vGPU placement neural networks; and utilizing the combined neural network selector to select a particular workload of at least one candidate workload, and execute the particular workload in a particular computing environment using a particular vGPU-enabled GPU.
9 . The method of claim 8 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a workload type of the particular workload.
10 . The method of claim 8 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a datacenter configuration of the particular computing environment.
11 . The method of claim 8 , wherein monitoring the workloads comprises executing a plurality of workloads comprising a plurality of different sets of a plurality of workload parameters, wherein the plurality of workloads are executed using the at least one computing environment configured according to the plurality of different datacenter configurations.
12 . The method of claim 11 , wherein a particular parameter of the plurality of workload parameters comprises: a measure of the composite efficiency metric calculated for a workload scaled by a geometric mean of the composite efficiency metric calculated for a respective one of a plurality of workloads currently in an arrival queue awaiting selection for placement.
13 . The method of claim 11 , wherein a particular set of of workload parameters comprises at least one of: a workload type that indicates a purpose or activity performed, a minimum graphics memory, a graphics memory requested as a fraction of a total GPU memory in the at least one computing environment, or any combination thereof.
14 . The method of claim 11 , wherein the at least one computing environment comprises a cluster of computing devices that provides host resources comprising a plurality of vGPU-enabled GPUs and at least one data store, wherein the at least one data store includes: a plurality of workloads, a scheduling service, and a plurality of vGPU placement neural networks.
15 . A non-transitory computer-readable medium comprising machine readable instructions, wherein the instructions, when executed by at least one processor, cause at least one computing device to at least:
identify workload data for a plurality of different datacenter configurations comprising at least one of: a plurality of different graphics processing unit (GPU) counts, a plurality of different virtual GPU (vGPU) scheduling policies, a plurality of different workload arrival rates, or any combination thereof, wherein the workload data is generated by monitoring workloads executed or simulated using at least one computing environment configured according to the plurality of different datacenter configurations; train a plurality of vGPU placement neural networks to maximize a composite efficiency metric based at least in part on the workload data identified for the plurality of different datacenter configurations, a respective one of the vGPU placement neural networks comprising at least two sets of layers; generate a combined neural network selector based on the plurality of vGPU placement neural networks; and utilize the combined neural network selector to select a particular workload of at least one candidate workload, and execute the particular workload in a particular computing environment using a particular vGPU-enabled GPU.
16 . The non-transitory computer-readable medium of claim 15 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a workload type of the particular workload.
17 . The non-transitory computer-readable medium of claim 15 , wherein the combined neural network selector selects the particular workload of the at least one candidate workload based at least in part on a datacenter configuration of the particular computing environment.
18 . The non-transitory computer-readable medium of claim 15 , wherein monitoring the workloads comprises executing a plurality of workloads comprising a plurality of different sets of a plurality of workload parameters, wherein the plurality of workloads are executed using the at least one computing environment configured according to the plurality of different datacenter configurations.
19 . The non-transitory computer-readable medium of claim 18 , wherein a particular parameter of the plurality of workload parameters comprises: a measure of the composite efficiency metric calculated for a workload scaled by a geometric mean of the composite efficiency metric calculated for a respective one of a plurality of workloads currently in an arrival queue awaiting selection for placement.
20 . The non-transitory computer-readable medium of claim 18 , wherein a particular parameter of the plurality of workload parameters comprises: a workload type that indicates a purpose or activity performed.Join the waitlist — get patent alerts
Track US2024036937A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.