US2021350233A1PendingUtilityA1

System and Method for Automated Precision Configuration for Deep Neural Networks

Assignee: DEEPLITE INCPriority: Nov 19, 2018Filed: Nov 18, 2019Published: Nov 11, 2021
Est. expiryNov 19, 2038(~12.3 yrs left)· nominal 20-yr term from priority
G06N 3/088G06N 3/08G06N 3/045G06F 18/2115G06N 3/082G06N 3/0985G06N 3/092G06N 3/0495G06N 3/0464G06N 3/063G06N 3/006G06N 3/0454G06K 9/6231
30
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

There is provided a system and method of automated precision configuration for deep neural networks. The method includes obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network; learning an optimal low-precision configuration of the architecture using constraints, the training data set, and the validation data set; and deploying the optimal configuration on the target device or process for use in the application.

Claims

exact text as granted — not AI-modified
1 . A method of automated precision configuration for deep neural networks, the method comprising:
 obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network;   learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and   deploying the optimal configuration on the target device or process for use in the application.   
     
     
         2 . The method of  claim 1 , wherein the optimal configuration is learned using a policy to generate an optimized model from the input model. 
     
     
         3 . The method of  claim 2 , wherein the optimal low-precision configuration of the optimal architecture is learned using the policy to generate a quantized network, the method further comprising:
 fine tuning the quantized network with a knowledge distillation process;   evaluating the fine-tuned network;   applying a reward function; and   iterating for at least one additional quantized network and selecting the optimal low-precision configuration.   
     
     
         4 . The method of  claim 3 , wherein selecting the optimal low-precision configuration comprises selecting a precision configuration that achieves the best reward as determined by the reward function, for the constraints on the target device or process. 
     
     
         5 . The method of  claim 1 , wherein learning the optimal low-precision configuration comprises exploiting low precision weights using reinforcement learning to learn the optimal low-precision configuration across the deep neural network. 
     
     
         6 . The method of  claim 5 , wherein each layer comprises a different precision. 
     
     
         7 . The method of  claim 1 , wherein the constraints comprise at least one of: accuracy, power, cost, supported precision, speed. 
     
     
         8 . The method of  claim 7 , wherein a computation constraint comprises a bit budget. 
     
     
         9 . The method of  claim 1 , wherein the application is an artificial intelligence-based application. 
     
     
         10 . A non-transitory computer readable medium comprising computer executable instructions for automated design space exploration for deep neural networks, the computer executable instructions comprising instructions for:
 obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network;   learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and   deploying the optimal configuration on the target device or process for use in the application.   
     
     
         11 . A deep neural network optimization engine configured to perform automated design space exploration for deep neural networks, the engine comprising a processor and memory, the memory comprising computer executable instructions for:
 obtaining an input model and one or more constraints associated with an application and/or target device or process used in the application configured to utilize a deep neural network;   learning an optimal low-precision configuration of the optimal architecture using the input model, constraints, a training data set, and a validation data set; and   deploying the optimal configuration on the target device or process for use in the application.   
     
     
         12 . The engine of  claim 11 , wherein the optimal configuration is learned using a policy to generate an optimized model from the input model. 
     
     
         13 . The engine of  claim 2 , wherein the optimal low-precision configuration of the optimal architecture is learned using the policy to generate a quantized network, further comprising instructions for:
 fine tuning the quantized network with a knowledge distillation process;   evaluating the fine-tuned network;   applying a reward function; and   iterating for at least one additional quantized network and selecting the optimal low-precision configuration.   
     
     
         14 . The engine of  claim 13 , wherein selecting the optimal low-precision configuration comprises selecting a precision configuration that achieves the best reward as determined by the reward function, for the constraints on the target device or process. 
     
     
         15 . The engine of  claim 11 , wherein learning the optimal low-precision configuration comprises exploiting low precision weights using reinforcement learning to learn the optimal low-precision configuration across the deep neural network. 
     
     
         16 . The engine of  claim 15 , wherein each layer comprises a different precision. 
     
     
         17 . The engine of  claim 11 , wherein the constraints comprise at least one of: accuracy, power, cost, supported precision, speed. 
     
     
         18 . The engine of  claim 17 , wherein a computation constraint comprises a bit budget. 
     
     
         19 . The engine of  claim 11 , wherein the application is an artificial intelligence-based application.

Join the waitlist — get patent alerts

Track US2021350233A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.