US2021350233A1PendingUtilityA1
System and Method for Automated Precision Configuration for Deep Neural Networks
Est. expiryNov 19, 2038(~12.3 yrs left)· nominal 20-yr term from priority
G06N 3/088G06N 3/08G06N 3/045G06F 18/2115G06N 3/082G06N 3/0985G06N 3/092G06N 3/0495G06N 3/0464G06N 3/063G06N 3/006G06N 3/0454G06K 9/6231
30
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
There is provided a system and method of automated precision configuration for deep neural networks. The method includes obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network; learning an optimal low-precision configuration of the architecture using constraints, the training data set, and the validation data set; and deploying the optimal configuration on the target device or process for use in the application.
Claims
exact text as granted — not AI-modified1 . A method of automated precision configuration for deep neural networks, the method comprising:
obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network; learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and deploying the optimal configuration on the target device or process for use in the application.
2 . The method of claim 1 , wherein the optimal configuration is learned using a policy to generate an optimized model from the input model.
3 . The method of claim 2 , wherein the optimal low-precision configuration of the optimal architecture is learned using the policy to generate a quantized network, the method further comprising:
fine tuning the quantized network with a knowledge distillation process; evaluating the fine-tuned network; applying a reward function; and iterating for at least one additional quantized network and selecting the optimal low-precision configuration.
4 . The method of claim 3 , wherein selecting the optimal low-precision configuration comprises selecting a precision configuration that achieves the best reward as determined by the reward function, for the constraints on the target device or process.
5 . The method of claim 1 , wherein learning the optimal low-precision configuration comprises exploiting low precision weights using reinforcement learning to learn the optimal low-precision configuration across the deep neural network.
6 . The method of claim 5 , wherein each layer comprises a different precision.
7 . The method of claim 1 , wherein the constraints comprise at least one of: accuracy, power, cost, supported precision, speed.
8 . The method of claim 7 , wherein a computation constraint comprises a bit budget.
9 . The method of claim 1 , wherein the application is an artificial intelligence-based application.
10 . A non-transitory computer readable medium comprising computer executable instructions for automated design space exploration for deep neural networks, the computer executable instructions comprising instructions for:
obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network; learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and deploying the optimal configuration on the target device or process for use in the application.
11 . A deep neural network optimization engine configured to perform automated design space exploration for deep neural networks, the engine comprising a processor and memory, the memory comprising computer executable instructions for:
obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network; learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and deploying the optimal configuration on the target device or process for use in the application.
12 . The engine of claim 11 , wherein the optimal configuration is learned using a policy to generate an optimized model from the input model.
13 . The engine of claim 2 , wherein the optimal low-precision configuration of the optimal architecture is learned using the policy to generate a quantized network, further comprising instructions for:
fine tuning the quantized network with a knowledge distillation process; evaluating the fine-tuned network; applying a reward function; and iterating for at least one additional quantized network and selecting the optimal low-precision configuration.
14 . The engine of claim 13 , wherein selecting the optimal low-precision configuration comprises selecting a precision configuration that achieves the best reward as determined by the reward function, for the constraints on the target device or process.
15 . The engine of claim 11 , wherein learning the optimal low-precision configuration comprises exploiting low precision weights using reinforcement learning to learn the optimal low-precision configuration across the deep neural network.
16 . The engine of claim 15 , wherein each layer comprises a different precision.
17 . The engine of claim 11 , wherein the constraints comprise at least one of: accuracy, power, cost, supported precision, speed.
18 . The engine of claim 17 , wherein a computation constraint comprises a bit budget.
19 . The engine of claim 11 , wherein the application is an artificial intelligence-based application.Join the waitlist — get patent alerts
Track US2021350233A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.