Neural network-based load balancing in distributed storage systems
Abstract
A system, method, and computer readable medium train machine learning model using server parameters corresponding to a plurality of servers. The training produces an optimized feature set of the machine learning model. The system, method, and computer readable medium assign a server class to a received request based on the optimized feature set. The system, method, and computer readable medium compute, based on the optimized feature set, estimate response times for one or more servers from the plurality of servers corresponding to the server class. The system, method, and computer readable medium forward the request to one of the one or more servers based on the estimate response times.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
training a machine learning model using server parameters corresponding to a plurality of servers, wherein the training produces an optimized feature set of the machine learning model; responsive to receiving a request, assigning a server class to the request based on the optimized feature set; computing, by a processing device and based on the optimized feature set, estimate response times for one or more servers from the plurality of servers corresponding to the server class; and forwarding the request to one of the one or more servers based on the estimate response times.
2 . The method of claim 1 , further comprising:
capturing performance parameters of the server that was forwarded the request while the server is processing the request; and retraining the machine learning model using the performance parameters.
3 . The method of claim 1 , further comprising:
computing servicing requests times of each of the one or more servers; and prioritizing the request based on the servicing requests times.
4 . The method of claim 1 , further comprising:
computing one or more current server loads for each of the one or more servers; identifying a number of static requests currently being processed by each one of the one or more servers; identifying a number of dynamic requests currently being processed by each one of the one or more servers; and computing the estimate response times for the one or servers based on their corresponding server load, the number of static requests currently being processed, and the number of dynamic requests currently being processed.
5 . The method of claim 1 , further comprising:
determining a request type of the request, wherein the request type is selected from the group consisting of a static request type and a dynamic request type; and assigning the server class to the request based on the request type.
6 . The method of claim 1 , further comprising:
analyzing computing resources dedicated to each of the plurality of servers; and assigning a server class to each of the plurality of servers based on their corresponding computing resources.
7 . The method of claim 1 , wherein the machine learning model comprises a classifier layer, a calculator layer, a decision layer, a forwarder layer, and a plurality of executor layers, wherein each one of the plurality of executor layers is assigned to one of the plurality of servers.
8 . A system comprising:
a memory; and a processing device operatively coupled to the memory, the processing device to:
train a machine learning model using server parameters corresponding to a plurality of servers, wherein the training produces an optimized feature set of the machine learning model;
responsive to receiving a request, assign a server class to the request based on the optimized feature set;
compute, based on the optimized feature set, estimate response times for one or more servers from the plurality of servers corresponding to the server class; and
forward the request to one of the one or more servers based on the estimate response times.
9 . The system of claim 8 , wherein the processing device is to:
capture performance parameters of the server that was forwarded the request while the server is processing the request; and retrain the machine learning model using the performance parameters.
10 . The system of claim 8 , wherein the processing device is to:
compute servicing requests times of each of the one or more servers; and prioritize the request based on the servicing requests times.
11 . The system of claim 8 , wherein the processing device is to:
compute one or more current server loads for each of the one or more servers; identify a number of static requests currently being processed by each one of the one or more servers; identify a number of dynamic requests currently being processed by each one of the one or more servers; and compute the estimate response times for the one or servers based on their corresponding server load, the number of static requests currently being processed, and the number of dynamic requests currently being processed.
12 . The system of claim 8 , wherein the processing device is to:
determine a request type of the request, wherein the request type is selected from the group consisting of a static request type and a dynamic request type; and assign the server class to the request based on the request type.
13 . The system of claim 8 , wherein the processing device is to:
analyze computing resources dedicated to each of the plurality of servers; and assign a server class to each of the plurality of servers based on their corresponding computing resources.
14 . The system of claim 8 , wherein the machine learning model comprises a classifier layer, a calculator layer, a decision layer, a forwarder layer, and a plurality of executor layers, wherein each one of the plurality of executor layers is assigned to one of the plurality of servers.
15 . A non-transitory computer readable medium, having instructions stored thereon which, when executed by a processing device, cause the processing device to:
train a machine learning model using server parameters corresponding to a plurality of servers, wherein the training produces an optimized feature set of the machine learning model; responsive to receiving a request, assign a server class to the request based on the optimized feature set; compute, by the processing device and based on the optimized feature set, estimate response times for one or more servers from the plurality of servers corresponding to the server class; and forward the request to one of the one or more servers based on the estimate response times.
16 . The non-transitory computer readable medium of claim 15 , wherein the processing device is to:
capture performance parameters of the server that was forwarded the request while the server is processing the request; and retrain the machine learning model using the performance parameters.
17 . The non-transitory computer readable medium of claim 15 , wherein the processing device is to:
compute servicing requests times of each of the one or more servers; and prioritize the request based on the servicing requests times.
18 . The non-transitory computer readable medium of claim 15 , wherein the processing device is to:
compute one or more current server loads for each of the one or more servers; identify a number of static requests currently being processed by each one of the one or more servers; identify a number of dynamic requests currently being processed by each one of the one or more servers; and compute the estimate response times for the one or servers based on their corresponding server load, the number of static requests currently being processed, and the number of dynamic requests currently being processed.
19 . The non-transitory computer readable medium of claim 15 , wherein the processing device is to:
determine a request type of the request, wherein the request type is selected from the group consisting of a static request type and a dynamic request type; and assign the server class to the request based on the request type.
20 . The non-transitory computer readable medium of claim 15 , wherein the processing device is to:
analyze computing resources dedicated to each of the plurality of servers; and assign a server class to each of the plurality of servers based on their corresponding computing resources.Join the waitlist — get patent alerts
Track US2024177050A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.