Method and apparatus for multi-level stepwise quantization for neural network
Abstract
A method and apparatus for multi-level stepwise quantization for neural network are provided. The apparatus sets a reference level by selecting a value from among values of parameters of the neural network in a direction from a high value equal to or greater than a predetermined value to a lower value, and performs learning based on the reference level. The setting of a reference level and the performing of learning are iteratively performed until the result of the reference level learning satisfies a predetermined value and there is no variable parameter that is updated during learning among the parameters.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A quantization method in a neural network, comprising:
setting a reference level by selecting a value from among values of parameters of the neural network in a direction from a high value equal to or greater than a predetermined value to a lower value; and performing reference level learning while the set reference level is fixed, wherein the setting of a reference level and the performing of reference level learning are iteratively performed until the result of the reference level learning satisfies a predetermined value and there is no variable parameter that is updated during learning among the parameters.
2 . The quantization method of claim 1 , further comprising:
when the result of the reference level learning does not satisfy the predetermined value, adding an offset level for the reference level and then performing offset level learning in which learning is performed while the offset level is fixed.
3 . The quantization method of claim 2 , wherein
the setting of a reference level, the performing of reference level learning, and the performing of offset level learning are iteratively performed until the result of the reference level learning or the result of the offset level learning satisfies a predetermined value and there is no variable parameter that is updated during learning among the parameters.
4 . The quantization method of claim 2 , wherein
the being fixed represents that no update to a parameter is performed during learning.
5 . The quantization method of claim 4 , wherein
the being fixed includes that parameters included in a setting range around the reference level or the offset level are fixed, and parameters not included in the setting range are variable parameters that are updated during learning.
6 . The quantization method of claim 2 , wherein
in the performing of offset level learning, the offset level is a level corresponding to a lowest value among parameters included in a set range around the reference level.
7 . The quantization method of claim 6 , wherein
the addition of the offset level is performed in a direction in which a scale is increased by a set multiple starting from a level corresponding to the lowest value.
8 . The quantization method of claim 2 , further comprising:
when the result of the reference level learning or the result of the offset level learning satisfies the predetermined value and there is no variable parameter that is updated during learning among the parameters, determining a quantization bit based on the reference level set so far and the offset level added so far.
9 . The quantization method of claim 8 , wherein
the determining of a quantization bit comprises: determining a quantization bit of parameters corresponding to the reference levels set so far according to a number of reference levels set so far; and determining a quantization bit of parameters corresponding to the offset levels added so far according to a number of offset levels added so far.
10 . The quantization method of claim 8 , further comprising:
before the determining of a quantization bit, setting remaining parameters to 0 except for parameters corresponding to the reference levels set so far and parameters corresponding to the offset levels added so far.
11 . The quantization method of claim 1 , wherein
the setting of a reference level comprises setting a maximum value among values of the parameters as a reference level, and then setting a random value in a direction from the maximum value to a minimum value.
12 . A quantization apparatus in a neural network, comprising:
an input interface device; and a processor configured to perform multi-level stepwise quantization for the neural network based on data input through the interface device, wherein the processor is configured to set a reference level by selecting a value from among values of parameters of the neural network in a direction from a high value equal to or greater than a predetermined value to a lower value, and perform learning based on the reference level, wherein the setting of a reference level and the performing of learning are iteratively performed until the result of the reference level learning satisfies a predetermined value and there is no variable parameter that is updated during learning among the parameters.
13 . The quantization apparatus of claim 12 , wherein
the processor is configured to perform the following operations: setting a reference level by selecting a value from among values of parameters of the neural network; performing reference level learning while the set reference level is fixed; and when the result of the reference level learning does not satisfy the predetermined value, adding an offset level for the reference level and then performing offset level learning in which learning is performed while the offset level is fixed, and wherein the setting of a reference level, the performing of reference level learning, and the performing of offset level learning are iteratively performed until the result of the reference level learning or the result of the offset level learning satisfies a predetermined value and there is no variable parameter that is updated during learning among the parameters.
14 . The quantization apparatus of claim 13 , wherein
the being fixed represents that no update to a parameter is performed during learning.
15 . The quantization apparatus of claim 14 , wherein
the being fixed includes that parameters included in a setting range around the reference level or the offset level are fixed, and parameters not included in the setting range are variable parameters that are updated during learning.
16 . The quantization apparatus of claim 13 , wherein
in the performing of offset level learning, the offset level is a level corresponding to a lowest value among parameters included in a set range around the reference level.
17 . The quantization apparatus of claim 16 , wherein
the addition of the offset level is performed in a direction in which a scale is increased by a set multiple starting from a level corresponding to the lowest value.
18 . The quantization apparatus of claim 13 , wherein
the processor is further configured to perform the following operation: when the result of the reference level learning or the result of the offset level learning satisfies the predetermined value and there is no variable parameter that is updated during learning among the parameters, determining a quantization bit based on the reference level set so far and the offset level added so far.
19 . The quantization apparatus of claim 13 , wherein
when performing the determining of a quantization bit, the processor is specifically configured to perform the following operation: determining a quantization bit of parameters corresponding to the reference levels set so far according to a number of reference levels set so far; and determining a quantization bit of parameters corresponding to the offset levels added so far according to a number of offset levels added so far.
20 . The quantization apparatus of claim 18 , wherein
before the determining of a quantization bit, the processor is further configured to perform the following operation: setting remaining parameters to 0 except for parameters corresponding to the reference levels set so far and parameters corresponding to the offset levels added so far.Join the waitlist — get patent alerts
Track US2021357753A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.