US2024256843A1PendingUtilityA1

Electronic apparatus for quantizing neural network model and control method thereof

Assignee: SAMSUNG ELECTRONICS CO LTDPriority: Jan 30, 2023Filed: Nov 27, 2023Published: Aug 1, 2024
Est. expiryJan 30, 2043(~16.5 yrs left)· nominal 20-yr term from priority
Inventors:Hyukjin Jeong
G06N 3/08G06N 3/045G06N 3/0495
42
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An electronic apparatus is provided. The electronic apparatus includes a memory, at least one processor connected to the memory and configured to control the electronic apparatus, and the processor may obtain a first neural network model comprising at least one layer that may be quantized, obtain test data used as an input of the first neural network model, obtain feature map data from each of at least one layer included in the first neural network model by inputting the test data to the first neural network model, obtain information about at least one of scaling or shifting to equalize channel-wise data from the feature map data obtained from each of the at least one layer, obtain a second neural network model in which channel-wise data of the feature map data is equalized by updating each of the at least one layer based on the obtained information, and obtain a quantized third neural network model corresponding to the first neural network model by quantizing the second neural network model based on the test data.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An electronic apparatus comprising:
 a memory; and   at least one processor connected to the memory and configured to control the electronic apparatus;   wherein the at least one processor is configured to:
 obtain a first neural network model comprising at least one layer that may be quantized, 
 obtain test data used as an input of the first neural network model, 
 obtain feature map data from each of at least one layer included in the first neural network model by inputting the test data to the first neural network model, 
 obtain information about at least one of scaling or shifting to equalize channel-wise data from the feature map data obtained from each of the at least one layer, 
 obtain a second neural network model in which channel-wise data of the feature map data is equalized by updating each of the at least one layer based on the obtained information, and 
 obtain a quantized third neural network model corresponding to the first neural network model by quantizing the second neural network model based on the test data. 
   
     
     
         2 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to obtain information about at least one of the scaling or the shifting based on a type of a preceding layer and a succeeding layer of the feature map data obtained from each of the at least one layer. 
     
     
         3 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to: in a case of obtaining the information about the scaling, based on scaling the feature map data of a first layer among the at least one layer:
 update the first layer by multiplying channel-wise scaling data by a weight and a bias of the first layer for each channel, and   update a second layer by dividing the second layer that is an immediately succeeding layer of the first layer by the channel-wise scaling data for each channel.   
     
     
         4 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to: in a case of obtaining the information about the shifting, based on shifting the feature map data of a third layer among the at least one layer:
 update the third layer by adding the channel-wise shifting data to a bias of the third layer by channels, and   update a fourth layer by multiplying the channel-wise shifting data by a weight of a fourth layer immediately succeeding the third layer by channels, adding the multiplication result by channels, and subtracting the channel-wise addition result from a bias of the fourth layer.   
     
     
         5 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to, in a case of obtaining the information about the scaling and the shifting, based on scaling the feature map data of a fifth layer among the at least one layer:
 apply scaling to the feature map data of the fifth layer, and   apply shifting to the feature map data to which the scaling is applied.   
     
     
         6 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to obtain information about at least one of the scaling or the shifting based on a channel having a maximum range among the feature map data obtained from each of the at least one layer. 
     
     
         7 . The electronic apparatus of  claim 6 , wherein the at least one processor is further configured to, in a case of obtaining the information about the scaling and the shifting, for each of the at least one layer:
 shift a range of a channel having the maximum range, and   scale or shift a range of remaining channels based on the shifted range.   
     
     
         8 . The electronic apparatus of  claim 6 , wherein the at least one processor is further configured to, in a case of obtaining the information about the scaling, based on scaling the feature map data of a first layer among the at least one layer, obtain information about the scaling so that a range of remaining channels other than a channel having the maximum range, among the feature map data of the first layer, is smaller by a preset value or more than a range of a channel having the maximum range among the feature map data of the first layer. 
     
     
         9 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to obtain the third neural network model by:
 quantizing at least one layer included in the second neural network model by channels, and   quantizing feature map data of each of the at least one layer included in the second neural network model by feature map data.   
     
     
         10 . The electronic apparatus of  claim 1 , wherein the at least one processor is further configured to quantize the second neural network model through affine transformation. 
     
     
         11 . A method of controlling an electronic apparatus, the method comprising:
 obtaining a first neural network model comprising at least one layer that may be quantized;   obtaining test data used as an input of the first neural network model;   obtaining feature map data from each of at least one layer included in the first neural network model by inputting the test data to the first neural network model;   obtaining information about at least one of scaling or shifting to equalize channel-wise data from the feature map data obtained from each of the at least one layer;   obtaining a second neural network model in which channel-wise data of the feature map data is equalized by updating each of the at least one layer based on the obtained information; and   obtaining a quantized third neural network model corresponding to the first neural network model by quantizing the second neural network model based on the test data.   
     
     
         12 . The method of  claim 11 , wherein the obtaining of information about at least one of scaling or shifting comprises obtaining information about at least one of the scaling or the shifting based on a type of a preceding layer and a succeeding layer of the feature map data obtained from each of the at least one layer. 
     
     
         13 . The method of  claim 11 , wherein the obtaining the second neural network model comprises, in a case of obtaining the information about the scaling, based on scaling the feature map data of a first layer among the at least one layer:
 updating the first layer by multiplying channel-wise scaling data by a weight and a bias of the first layer for each channel; and   updating a second layer by dividing the second layer that is an immediately succeeding layer of the first layer by the channel-wise scaling data for each channel.   
     
     
         14 . The method of  claim 11 , wherein the obtaining the second neural network model comprises, in a case of obtaining the information about the shifting, based on shifting the feature map data of a third layer among the at least one layer:
 updating the third layer by adding the channel-wise shifting data to a bias of the third layer by channels; and   updating a fourth layer by multiplying the channel-wise shifting data by a weight of a fourth layer immediately succeeding the third layer by channels, adding the multiplication result by channels, and subtracting the channel-wise addition result from a bias of the fourth layer.   
     
     
         15 . The method of  claim 11 , wherein the obtaining the second neural network model comprises, in a case of obtaining the information about the scaling and the shifting, based on scaling the feature map data of a fifth layer among the at least one layer:
 applying scaling to the feature map data of the fifth layer; and   applying shifting to the feature map data to which the scaling is applied.   
     
     
         16 . The method of  claim 11 , wherein the obtaining of the information about at least one of scaling or shifting comprises detecting an equalization pattern from the obtained feature map data. 
     
     
         17 . The method of  claim 16 ,
 wherein the at least one layer comprises a plurality of layers, and   wherein the detecting of the equalization pattern from the obtained feature map data comprises:
 comparing each of the plurality of layers with an adjacent layer among the plurality of layers; and 
 determining the equalization pattern based on whether the compared layers are of a convolution layer type, a deconvolution layer type, or a combination of the convolution layer type and the deconvolution layer type.

Join the waitlist — get patent alerts

Track US2024256843A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.