Apparatus and method for split processing of model
Abstract
An apparatus and method for split processing of a model are provided. The apparatus for the split processing of the model includes a memory including instructions and a processor electrically connected to the memory and configured to execute the instructions. When the instructions are executed by the processor, the processor may be configured to perform a plurality of operations. The plurality of operations may include obtaining information on a plurality of computing nodes that uses at least one layer among a plurality of layers of a model for an artificial intelligence (AI)-based service, obtaining a requirement for the AI-based service, and controlling split processing of the model based on the information and the requirement.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An apparatus for split processing of a model, the apparatus comprising:
a memory comprising instructions; and a processor electrically connected to the memory and configured to execute the instructions, wherein, when the instructions are executed by the processor, the processor is configured to perform a plurality of operations, wherein the plurality of operations comprises: obtaining information on a plurality of computing nodes that uses at least one layer among a plurality of layers of a model for an artificial intelligence (AI)-based service; obtaining a requirement for the AI-based service; and controlling split processing of the model based on the information and the requirement.
2 . The apparatus of claim 1 , wherein the obtaining of the information comprises receiving at least one of first information on computing of the plurality of computing nodes or second information on a state of the plurality of computing nodes from the plurality of computing nodes.
3 . The apparatus of claim 2 , wherein the second information comprises information on mobility of the plurality of computing nodes.
4 . The apparatus of claim 1 , wherein the requirement comprises at least one of computing latency or computing accuracy of the plurality of computing nodes that are required for the AI-based service.
5 . The apparatus of claim 4 , wherein the computing latency comprises at least one of computing latency of the plurality of computing nodes in a learning process or computing latency of the plurality of computing nodes in an inference process.
6 . The apparatus of claim 4 , wherein the computing accuracy comprises at least one of computing accuracy of the plurality of computing nodes in a learning process or computing accuracy of the plurality of computing nodes in an inference process.
7 . The apparatus of claim 1 , wherein the controlling of the split processing of the model comprises determining a split point for the plurality of layers.
8 . The apparatus of claim 1 , wherein the controlling of the split processing of the model comprises transmitting data related to a first computing node that is included in the plurality of computing nodes to a second computing node that is not included in the plurality of computing nodes.
9 . The apparatus of claim 8 , wherein the transmitting of the data comprises transmitting data on the model of the first computing node to the second computing node.
10 . The apparatus of claim 9 , wherein the transmitting of the data on the model to the second computing node comprises:
requesting data on the model from the first computing node based on information related to at least one computing node other than the first computing node among the plurality of computing nodes; and transmitting the data on the model to the second computing node.
11 . A method for split processing of a model, the method comprising:
obtaining information on a plurality of computing nodes that uses at least one layer among a plurality of layers of a model for an artificial intelligence (AI)-based service; obtaining a requirement for the AI-based service; and controlling split processing of the model based on the information and the requirement.
12 . The method of claim 11 , wherein the obtaining of the information comprises receiving at least one of first information on computing of the plurality of computing nodes or second information on a state of the plurality of computing nodes from the plurality of computing nodes.
13 . The method of claim 12 , wherein the second information comprises information on mobility of the plurality of computing nodes.
14 . The method of claim 11 , wherein the requirement comprises at least one of computing latency or computing accuracy of the plurality of computing nodes that are required for the AI-based service.
15 . The method of claim 14 , wherein the computing latency comprises at least one of computing latency of the plurality of computing nodes in a learning process or computing latency of the plurality of computing nodes in an inference process.
16 . The method of claim 14 , wherein the computing accuracy comprises at least one of computing accuracy of the plurality of computing nodes in a learning process or computing accuracy of the plurality of computing nodes in an inference process.
17 . The method of claim 11 , wherein the controlling of the split processing of the model comprises determining a split point for the plurality of layers.
18 . The method of claim 11 , wherein the controlling of the split processing of the model comprises transmitting data related to a first computing node that is included in the plurality of computing nodes to a second computing node that is not included in the plurality of computing nodes.
19 . The method of claim 18 , wherein the transmitting of the data comprises transmitting data on the model of the first computing node to the second computing node.
20 . The method of claim 19 , wherein the transmitting of the data on the model to the second computing node comprises:
requesting data on the model from the first computing node based on information related to at least one computing node other than the first computing node among the plurality of computing nodes; and transmitting the data on the model to the second computing node.Join the waitlist — get patent alerts
Track US2024185101A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.