Artificial neural network training using flexible floating point tensors
Abstract
Thus, the present disclosure is directed to systems and methods for training neural networks using a tensor that includes a plurality of FP16 values and a plurality of bits that define an exponent shared by some or all of the FP16 values included in the tensor. The FP16 values may include IEEE 754 format 16-bit floating point values and the tensor may include a plurality of bits defining the shared exponent. The tensor may include a shared exponent and FP16 values that include a variable bit-length mantissa and a variable bit-length exponent that may be dynamically set by processor circuitry. The tensor may include a shared exponent and FP16 values that include a variable bit-length mantissa; a variable bit-length exponent that may be dynamically set by processor circuitry; and a shared exponent switch set by the processor circuitry to selectively combine the FP16 value exponent with the shared exponent.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . An apparatus to train a machine learning model, the apparatus comprising:
processor circuitry; and a storage device accessible by the processor circuitry, the storage device including machine readable instructions to cause the processor circuitry to:
train the machine learning model using a tensor having a sequence of 16-bit floating point values, a part of the sequence of the 16-bit floating point values including respective mantissa values, at least two of the 16-bit floating point values associated with a shared exponent.
2 . The apparatus of claim 1 , wherein at least one of the 16-bit floating point values is not associated with the shared exponent.
3 . The apparatus of claim 1 , wherein a first one of the 16-bit floating point values includes a first sign bit.
4 . The apparatus of claim 1 , wherein the processor circuitry is to adjust a value of the shared exponent during training of the machine learning model.
5 . The apparatus of claim 4 , wherein the processor circuitry is to adjust the value of the shared exponent in response to a prediction of at least one of a future overflow or a future underflow.
6 . The apparatus of claim 1 , wherein the respective mantissa values are represented by six bits.
7 . The apparatus of claim 1 , wherein a first one of the 16-bit floating point values includes a bit representing whether the shared exponent is to be used.
8 . The apparatus of claim 1 , wherein the processor circuitry is to store the machine learning model in the storage device.
9 . At least one non-transitory computer readable storage medium comprising instructions that, when executed, cause at least one processor to at least:
train a machine learning model using a tensor having a sequence of 16-bit floating point values, a part of the sequence of the 16-bit floating point values including respective mantissa values, at least two of the 16-bit floating point values associated with a shared exponent; and store the trained machine learning model.
10 . The at least one non-transitory computer readable storage medium of claim 9 , wherein at least one of the 16-bit floating point values is not associated with the shared exponent.
11 . The at least one non-transitory computer readable storage medium of claim 9 , wherein a first one of the 16-bit floating point values includes a first sign bit.
12 . The at least one non-transitory computer readable storage medium of claim 9 , wherein the instructions, when executed, cause the at least one processor to adjust a value of the shared exponent during training of the machine learning model.
13 . The at least one non-transitory computer readable storage medium of claim 9 , wherein the respective mantissa values are represented by six bits.
14 . The at least one non-transitory computer readable storage medium of claim 9 , wherein a first one of the 16-bit floating point values includes a bit representing whether the shared exponent is to be used.
15 . A method comprising:
training a machine learning model using a tensor having a sequence of 16-bit floating point values, a part of the sequence of the 16-bit floating point values including respective mantissa values, at least two of the 16-bit floating point values associated with a shared exponent; and storing the trained machine learning model.
16 . The method of claim 15 , wherein at least one of the 16-bit floating point values is not associated with the shared exponent.
17 . The method of claim 15 , wherein a first one of the 16-bit floating point values includes a first sign bit.
18 . The method of claim 15 , further including adjusting a value of the shared exponent during training of the machine learning model.
19 . The method of claim 15 , wherein the respective mantissa values are represented by six bits.
20 . The method of claim 15 , wherein a first one of the 16-bit floating point values includes a bit representing whether the shared exponent is to be used.Join the waitlist — get patent alerts
Track US2024028905A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.