Artificial intelligence inferencing workload placement to minimize latency in a heterogeneous environment
Abstract
A method for managing inferencing workload placement based on latency minimization includes performing an initial workload placement of the inferencing workload to assign the inferencing workload to a first production environment of the plurality of production environments, after performing the initial workload placement, monitoring: execution of the inferencing workload on the first production environment, and communication between the first production environment and a front-end environment, to obtain telemetry data associated with the execution and the communication, performing a latency analysis using the telemetry data to generate a placement recommendation, making a determination that the placement recommendation specifies a second production environment of the plurality of production environments, and based on the determination, initiating deployment of the inferencing workload to the second production environment.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for managing workload placement, the method comprising:
obtaining, by a workload placement service, a request for assigning an inferencing workload to one of a plurality of production environments based on latency minimization; in response to the request:
performing an initial workload placement of the inferencing workload to assign the inferencing workload to a first production environment of the plurality of production environments;
after performing the initial workload placement, monitoring:
execution of the inferencing workload on the first production environment, and
communication between the first production environment and a front-end environment, to obtain telemetry data associated with the execution and the communication;
performing a latency analysis using the telemetry data to generate a placement recommendation;
making a determination that the placement recommendation specifies a second production environment of the plurality of production environments; and
based on the determination, initiating deployment of the inferencing workload to the second production environment.
2 . The method of claim 1 , wherein the inferencing workload comprises implementing a generative artificial intelligence (AI) model.
3 . The method of claim 2 , wherein the front-end environment comprises a front-end device operated by a user utilizing the generative AI model to obtain an inferencing payload.
4 . The method of claim 1 , wherein the latency analysis is further based on causal variables associated with latency in the communication.
5 . The method of claim 4 , wherein the causal variables comprise at least one of: clock speed of a graphics processing unit (GPU) of the first production environment, a number of GPUs used for the inferencing workload in the first production environment, a second number of GPUs available in the second production environment, and latency between GPUs executing the inferencing workload.
6 . The method of claim 1 , wherein the first production environment is a computing device of an on-premise environment.
7 . The method of claim 1 , wherein the first production environment is a computing device of a cloud environment operatively connected to the front-end environment via a wide area network.
8 . A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for managing information handling systems, the method comprising:
obtaining, by a workload placement service, a request for assigning an inferencing workload to one of a plurality of production environments based on latency minimization; in response to the request:
performing an initial workload placement of the inferencing workload to assign the inferencing workload to a first production environment of the plurality of production environments;
after performing the initial workload placement, monitoring:
execution of the inferencing workload on the first production environment, and
communication between the first production environment and a front-end environment, to obtain telemetry data associated with the execution and the communication;
performing a latency analysis using the telemetry data to generate a placement recommendation;
making a determination that the placement recommendation specifies a second production environment of the plurality of production environments; and
based on the determination, initiating deployment of the inferencing workload to the second production environment.
9 . The non-transitory computer readable medium of claim 8 , wherein the inferencing workload comprises implementing a generative artificial intelligence (AI) model.
10 . The non-transitory computer readable medium of claim 9 , wherein the front-end environment comprises a front-end device operated by a user utilizing the generative AI model to obtain an inferencing payload.
11 . The non-transitory computer readable medium of claim 8 , wherein the latency analysis is further based on causal variables associated with latency in the communication.
12 . The non-transitory computer readable medium of claim 11 , wherein the causal variables comprise at least one of: clock speed of a graphics processing unit (GPU) of the first production environment, a number of GPUs used for the inferencing workload in the first production environment, a second number of GPUs available in the second production environment, and latency between GPUs executing the inferencing workload.
13 . The non-transitory computer readable medium of claim 8 , wherein the first production environment is a computing device of an on-premise environment.
14 . The non-transitory computer readable medium of claim 8 , wherein the first production environment is a computing device of a cloud environment operatively connected to the front-end environment via a wide area network.
15 . A system, comprising:
a processor; and memory including instructions, which when executed by the processor, perform a method comprising:
obtaining, by a workload placement service, a request for assigning an inferencing workload to one of a plurality of production environments based on latency minimization;
in response to the request:
performing an initial workload placement of the inferencing workload to assign the inferencing workload to a first production environment of the plurality of production environments;
after performing the initial workload placement, monitoring:
execution of the inferencing workload on the first production environment, and
communication between the first production environment and a front-end environment, to obtain telemetry data associated with the execution and the communication;
performing a latency analysis using the telemetry data to generate a placement recommendation;
making a determination that the placement recommendation specifies a second production environment of the plurality of production environments; and
based on the determination, initiating deployment of the inferencing workload to the second production environment.
16 . The system of claim 15 , wherein the inferencing workload comprises implementing a generative artificial intelligence (AI) model.
17 . The system of claim 16 , wherein the front-end environment comprises a front-end device operated by a user utilizing the generative AI model to obtain an inferencing payload.
18 . The system of claim 15 , wherein the latency analysis is further based on causal variables associated with latency in the communication, and clock speed of a graphics processing unit (GPU) of the first production environment, a number of GPUs used for the inferencing workload in the first production environment, a second number of GPUs available in the second production environment, and latency between GPUs executing the inferencing workload.
19 . The system of claim 15 , wherein the first production environment is a computing device of an on-premise environment.
20 . The system of claim 15 , wherein the first production environment is a computing device of a cloud environment operatively connected to the front-end environment via a wide area network.Join the waitlist — get patent alerts
Track US2025238274A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.