Automated methods for conversions to a lower precision data format
Abstract
Aspects of the present invention are directed to computer-implemented techniques for performing data compression and conversion between data formats of varying degrees of precision, and more particularly for improving the inferencing (application) of artificial neural networks using a reduced precision (e.g., INT8) data format. Embodiments of the present invention generate candidate conversions of data output, then employ a relative measure of quality to identify the candidate conversion with the greatest accuracy (i.e., least divergence from the original higher precision values). The representation can be then be used during inference to perform computations on the resulting output data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A processor, comprising:
one or more circuits to cause a neural network to generate data having a first precision based, at least in part, on the neural network being trained using data having a second precision that is greater than the first precision.
2 . The processor of claim 1 , wherein the first precision corresponds to a selected conversion of a plurality of candidate conversions for the data having the second precision.
3 . The processor of claim 2 , wherein the one or more circuits are to generate the plurality of candidate conversions at least in part by referencing activation data for a layer of the neural network and creating a histogram of activation, the plurality of candidate conversions determined based on the histogram.
4 . The processor of claim 3 , wherein the histogram comprises a plurality of bins and the activation data is distributed across the plurality of bins, and wherein each conversion of the plurality of candidate conversions has a different saturation threshold.
5 . The processor of claim 2 , wherein the one or more circuits are to determine the selected conversion at least in part by determining a divergence for each conversion of the plurality of candidate conversions from a calibration data set, and selecting the saturation threshold corresponding to the conversion with the least divergence from the reference higher precision distribution.
6 . The processor of claim 5 , wherein determining the divergence comprises applying a metric for measuring directed divergence between the plurality of candidate conversions and the reference higher precision distribution.
7 . The processor of claim 6 , wherein the metric comprises determining a Kullback-Leibler divergence
8 . The processor of claim 2 , wherein the plurality of candidate conversions are expressed in a lower precision format, and wherein at least one of the calibration data set and activation data for a layer of the neural network is expressed in the higher precision format.
9 . The processor of claim 8 , wherein the plurality of candidate conversions comprise a plurality of quantized distributions of activations for the layer of the neural network that correspond to a range of values between zero and a maximum absolute value comprised in the activation data.
10 . A method, comprising:
causing a neural network to generate data having a first precision based, at least in part, on the neural network being trained using data having a second precision that is greater than the first precision.
11 . The method of claim 10 , wherein the first precision corresponds to a selected conversion of a plurality of candidate conversions for the data having the second precision.
12 . The method of claim 11 , further comprising:
generating the plurality of candidate conversions at least in part by referencing activation data for a layer of the neural network and creating a histogram of activation, the plurality of candidate conversions determined based on the histogram.
13 . The method of claim 12 , wherein the histogram comprises a plurality of bins and the activation data is distributed across the plurality of bins, and wherein each conversion of the plurality of candidate conversions has a different saturation threshold.
14 . The method of claim 11 , further comprising:
determining the selected conversion by determining a divergence for each conversion of the plurality of candidate conversions from a calibration data set, and selecting the saturation threshold corresponding to the conversion with the least divergence from the reference higher precision distribution.
15 . A system, comprising:
one or more processors to cause a neural network to generate data having a first precision based, at least in part, on the neural network being trained using data having a second precision that is greater than the first precision; and memory for storing network parameters for the neural network.
16 . The system of claim 15 , wherein the first precision corresponds to a selected conversion of a plurality of candidate conversions for the data having the second precision.
17 . The system of claim 16 , wherein the one or more processors are further to generate the plurality of candidate conversions at least in part by referencing activation data for a layer of the neural network and creating a histogram of activation, the plurality of candidate conversions determined based on the histogram.
18 . The system of claim 17 , wherein the histogram comprises a plurality of bins and the activation data is distributed across the plurality of bins, and wherein each conversion of the plurality of candidate conversions has a different saturation threshold.
19 . The system of claim 16 , wherein the one or more processors are further to determine the selected conversion by determining a divergence for each conversion of the plurality of candidate conversions from a calibration data set, and selecting the saturation threshold corresponding to the conversion with the least divergence from the reference higher precision distribution.
20 . The system of claim 16 , wherein the plurality of candidate conversions are expressed in a lower precision format, and wherein at least one of the calibration data set and activation data for a layer of the neural network is expressed in the higher precision format, and wherein the plurality of candidate conversions comprise a plurality of quantized distributions of activations for the layer of the neural network that correspond to a range of values between zero and a maximum absolute value comprised in the activation data.Join the waitlist — get patent alerts
Track US2021256348A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.