US2024012685A1PendingUtilityA1
Swarm multi-agent reinforcement learning-based pipeline for workload placement
Est. expiryJul 11, 2042(~15.9 yrs left)· nominal 20-yr term from priority
G06F 2209/501G06F 9/5088G06F 2209/506G06F 9/5011G06F 9/5027
47
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Multi-agent reinforcement learning-based workload placement is disclosed. A placement engine is configured to use the state of a system and actual rewards to generate expected rewards that correspond to actions. Agents can take actions for corresponding workloads based on the expected rewards output by the placement engine. This allows workloads to be placed in a manner that conservers power relative to load placement policies while helping avoid service level agreement violations.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
receiving input into a placement engine, the input including an actual reward of a workload operating in an environment included in resources and a state; generating expected rewards including a first expected reward and a second expected reward; performing the first action on the workload, by an agent associated with the workload, when the first expected reward is higher than the second expected reward; and performing the second action on the workload when the second expected reward is higher than the first expected reward.
2 . The method of claim 1 , wherein the environment comprises a current virtual machine and wherein the actual reward corresponds to a service level agreement metric of the workload operating at the current virtual machine.
3 . The method of claim 2 , wherein the state incudes a one hot encoding style of all environments, each of the environments including a virtual machine.
4 . The method of claim 3 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload, using floating point values.
5 . The method of claim 2 , wherein the first action is to keep the workload at the current virtual machine and wherein the second action is to migrate the workload to a different virtual machine.
6 . The method of claim 1 , wherein the placement engine comprises a neural network configured to map the input to expected rewards.
7 . The method of claim 1 , wherein the placement engine outputs an expected reward for performing an action relative to each of the virtual machine in the resources.
8 . The method of claim 1 , further comprising adjusting a reward function when an SLA violation is detected.
9 . The method of claim 8 , wherein the reward function is:
f
(
Δ
,
σ
L
,
σ
R
)
=
-
(
Δ
)
2
e
2
σ
L
2
if
Δ
>
1
,
otherwise
-
(
Δ
)
2
e
2
σ
L
2
-
1
,
wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.
10 . The method of claim 1 , wherein the placement engine is trained by randomly migrating workloads amongst virtual machines in the resources.
11 . The method of claim 1 , wherein the placement engine is configured to place the workload in a manner that includes both minimum virtual machine placement and load balancing.
12 . A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising:
receiving input into a placement engine, the input including an actual reward of a workload operating in an environment included in resources and a state;
generating expected rewards including a first expected reward and a second expected reward;
performing the first action on the workload, by an agent associated with the workload, when the first expected reward is higher than the second expected reward; and
performing the second action on the workload when the second expected reward is higher than the first expected reward.
13 . The non-transitory storage medium of claim 12 , wherein the environment comprises a current virtual machine and wherein the actual reward corresponds to a service level agreement metric of the workload operating at the current virtual machine.
14 . The non-transitory storage medium of claim 13 , wherein the state incudes a one hot encoding style of all environments, each of the environments including a virtual machine.
15 . The non-transitory storage medium of claim 14 , wherein the one hot encoding style includes a resource usage per virtual machine, a resource usage per workload, a state of each workload, and a time to completion for each workload using floating point values.
16 . The non-transitory storage medium of claim 13 , wherein the first action is to keep the workload at the current virtual machine and wherein the second action is to migrate the workload to a different virtual machine.
17 . The non-transitory storage medium of claim 12 , wherein the placement engine comprises a neural network configured to map the input to expected rewards, wherein the placement engine is configured to place the workload in a manner that includes both minimum virtual machine placement and load balancing.
18 . The non-transitory storage medium of claim 12 , wherein the placement engine outputs an expected reward for performing an action relative to each of the virtual machine in the resources.
19 . The non-transitory storage medium of claim 12 , further comprising adjusting a reward function when an SLA violation is detected.
20 . The non-transitory storage medium of claim 19 , wherein the reward function is:
f
(
Δ
,
σ
L
,
σ
R
)
=
-
(
Δ
)
2
e
2
σ
L
2
if
Δ
>
1
,
otherwise
-
(
Δ
)
2
e
2
σ
L
2
-
1
,
wherein Δ is a difference between an SLA response time metric and an actual response time for the workload in the environment, wherein σ L and σ R define, respectively, how fast a left and a right portion of the reward function decay.Join the waitlist — get patent alerts
Track US2024012685A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.