Virtualized computing resource management for machine learning model-based processing in computing environment
Abstract
Techniques are disclosed for virtualized computing resource management for machine learning model-based processing in a computing environment. For example, a method maintains one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed. After creation and performance of the one or more initializations, each of the one or more virtualized computing resources is placed in an idle state. The method then receives a machine learning model-based request, and removes at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
maintaining one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed, and further wherein each of the one or more virtualized computing resources is placed in an idle state after creation and performance of the one or more initializations; receiving a machine learning model-based request; and removing at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request; wherein the maintaining, receiving and removing steps are performed by at least one processor and at least one memory storing executable computer program instructions.
2 . The method of claim 1 , wherein the machine learning model-based request comprises an inference serving request.
3 . The method of claim 2 , wherein the at least one virtualized computing resource removed from the idle state is used to process the inference serving request by:
loading a trained machine learning model; processing input associated with the inference serving request using the trained machine learning model; and returning a result of the input processing by the trained machine learning model.
4 . The method of claim 1 , wherein the one or more initializations caused to be performed comprise initializing an operating system process.
5 . The method of claim 1 , wherein the one or more initializations caused to be performed comprise initializing a machine learning framework.
6 . The method of claim 1 , wherein the one or more initializations caused to be performed comprise initializing an accelerator.
7 . The method of claim 1 , wherein maintaining the one or more virtualized computing resources further comprises:
creating an additional virtualized computing resource and causing one or more initializations to be performed; and placing the additional virtualized computing resource in an idle state after the additional virtualized computing resource is created and the one or more initializations are performed.
8 . The method of claim 1 , wherein the one or more virtualized computing resources comprise one or more containers.
9 . The method of claim 8 , wherein the at least one processor and the at least one memory comprise a worker node in a container orchestration framework.
10 . The method of claim 9 , wherein the worker node is part of an edge computing platform.
11 . An apparatus, comprising:
at least one processor and at least one memory storing computer program instructions wherein, when the at least one processor executes the computer program instructions, the apparatus is configured to: maintain one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed, and further wherein each of the one or more virtualized computing resources is placed in an idle state after creation and performance of the one or more initializations; receive a machine learning model-based request; and remove at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request.
12 . The apparatus of claim 11 , wherein the machine learning model-based request comprises an inference serving request.
13 . The apparatus of claim 12 , wherein the at least one virtualized computing resource removed from the idle state is used to process the inference serving request by:
loading a trained machine learning model; processing input associated with the inference serving request using the trained machine learning model; and returning a result of the input processing by the trained machine learning model.
14 . The apparatus of claim 11 , wherein the one or more initializations caused to be performed comprise initializing one or more of an operating system process, a machine learning framework, and an accelerator.
15 . The apparatus of claim 11 , wherein the apparatus is further configured to maintain the one or more virtualized computing resources by:
creating an additional virtualized computing resource and causing one or more initializations to be performed; and placing the additional virtualized computing resource in an idle state after the additional virtualized computing resource is created and the one or more initializations are performed.
16 . The apparatus of claim 11 , wherein the one or more virtualized computing resources comprise one or more containers, the at least one processor and the at least one memory comprise a worker node in a container orchestration framework, and the worker node is part of an edge computing platform.
17 . A computer program product stored on a non-transitory computer-readable medium and comprising machine executable instructions, the machine executable instructions, when executed, causing a processing device to perform steps of:
maintaining one or more virtualized computing resources, wherein each of the one or more virtualized computing resources is created and one or more initializations are caused to be performed, and further wherein each of the one or more virtualized computing resources is placed in an idle state after creation and performance of the one or more initializations; receiving a machine learning model-based request; and removing at least one of the one or more virtualized computing resources from the idle state to process the machine learning model-based request.
18 . The computer program product of claim 17 , wherein the machine learning model-based request comprises an inference serving request.
19 . The computer program product of claim 17 , wherein the one or more initializations caused to be performed comprise initializing one or more of an operating system process, a machine learning framework, and an accelerator.
20 . The computer program product of claim 17 , wherein maintaining the one or more virtualized computing resources further comprises:
creating an additional virtualized computing resource and causing one or more initializations to be performed; and placing the additional virtualized computing resource in an idle state after the additional virtualized computing resource is created and the one or more initializations are performed.Join the waitlist — get patent alerts
Track US2023273837A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.