Numerical representation for neural networks
Abstract
Techniques in advanced deep learning provide improvements in one or more of accuracy, performance, and energy efficiency. An array of processing elements comprising a portion of a neural network accelerator performs flow-based computations on wavelets of data. Each processing element has a respective compute element and a respective routing element. Each compute element has a respective floating-point unit enabled to optionally and/or selectively perform floating-point operations in accordance with a programmable exponent bias and/or various floating-point computation variations. In some circumstances, the programmable exponent bias and/or the floating-point computation variations enable neural network processing with improved accuracy, decreased training time, decreased inference latency, and/or increased energy efficiency.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . An apparatus comprising:
a programmable processor enabled to execute instructions comprising floating-point instructions; exponent bias and exponent size control resources of the programmable processor enabled to maintain respective state representing a current exponent bias of a plurality of exponent biases and a current exponent size of a plurality of exponent sizes; a floating-point unit of the programmable processor enabled to perform floating-point operations corresponding to the floating-point instructions and in accordance with floating-point operands and floating-point results each comprising a respective biased exponent, the floating-point operations interpreting the biased exponents of the floating-point operands and producing the biased exponents of the floating-point results in accordance with the current exponent bias and the current exponent size; and wherein the instructions further comprise first and second load control resource instructions for respective operations to update the respective state maintained in the exponent bias and exponent size control resources.
3 . The apparatus of claim 2 , wherein:
the exponent size of the exponent size control resource specifies one of a plurality of floating-point formats, a first of the plurality of floating-point formats comprising a biased exponent of N bits and a mantissa of M bits, and a second of the plurality of floating-point formats comprising a biased exponent of (N+delta) bits and a mantissa of (M−delta) bits.
4 . The apparatus of claim 2 , wherein the programmable processor is one of a plurality of like programmable processors fabricated on a wafer in accordance with wafer-scale integration, and a datacenter element enabled to perform neural network processing comprises the wafer.
5 . The apparatus of claim 2 , wherein the programmable processor comprises one or more hardware registers comprising at least one of the exponent bias and the exponent size control resources, and at least one of the first and the second load control resource instructions is a load hardware register instruction.
6 . The apparatus of claim 2 , wherein at least one of the exponent bias and the exponent size control resources is memory-mapped, and at least one of the first and the second load control resource instructions is a memory store instruction.
7 . The apparatus of claim 2 , wherein:
the instructions further comprise a third load control resource instruction, the apparatus further comprises a rounding mode control resource of the programmable processor enabled to maintain state representing a current rounding mode of a plurality of rounding modes, at least one of the rounding modes comprising saturating the floating-point results to a value comprising a maximum biased exponent, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current rounding mode.
8 . The apparatus of claim 2 , wherein:
the instructions further comprise a third load control resource instruction, the apparatus further comprises a flush-to-zero control resource of the programmable processor enabled to maintain state representing a current flush-to-zero mode of a plurality of flush-to-zero modes, at least one of the flush-to-zero modes comprising flushing subnormal results to zero, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current flush-to-zero mode.
9 . An apparatus comprising:
a programmable processor enabled to execute instructions comprising floating-point instructions; exponent bias and maximum biased exponent normal control resources of the programmable processor enabled to maintain respective state representing a current exponent bias of a plurality of exponent biases and a current maximum biased exponent normal specifying one of a plurality of maximum biased exponent modes, the maximum biased exponent modes comprising a first mode using a maximum biased exponent to represent a biased exponent; a floating-point unit of the programmable processor enabled to perform floating-point operations corresponding to the floating-point instructions and in accordance with floating-point operands and floating-point results each comprising a respective biased exponent, the floating-point operations interpreting the biased exponents of the floating-point operands and producing the biased exponents of the floating-point results in accordance with the current exponent bias and the current maximum biased exponent normal; and wherein the instructions further comprise first and second load control resource instructions for respective operations to update the respective state maintained in the exponent bias and maximum biased exponent normal control resources.
10 . The apparatus of claim 9 , wherein:
the instructions further comprise a third load control resource instruction, the apparatus further comprises a rounding mode control resource of the programmable processor enabled to maintain state representing a current rounding mode of a plurality of rounding modes, at least one of the rounding modes comprising saturating the floating-point results to a value comprising a maximum biased exponent, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current rounding mode.
11 . The apparatus of claim 9 , wherein the maximum biased exponent modes further comprise a second mode using the maximum biased exponent to represent at least one of: infinite numbers and IEEE compatible NaN.
12 . The apparatus of claim 9 , wherein the programmable processor is one of a plurality of like programmable processors fabricated on a wafer in accordance with wafer-scale integration, and a datacenter element enabled to perform neural network processing comprises the wafer.
13 . The apparatus of claim 9 , wherein the programmable processor comprises one or more hardware registers comprising at least one of the exponent bias and the maximum biased exponent normal control resources, and at least one of the first and the second load control resource instructions is a load hardware register instruction.
14 . The apparatus of claim 9 , wherein at least one of the exponent bias and the maximum biased exponent normal control resources is memory-mapped, and at least one of the first and the second load control resource instructions is a memory store instruction.
15 . An apparatus comprising:
a programmable processor enabled to execute instructions comprising floating-point instructions; exponent bias and zero biased exponent mode control resources of the programmable processor enabled to maintain respective state representing a current exponent bias of a plurality of exponent biases and a zero biased exponent mode of a plurality of zero biased exponent modes, the zero biased exponent modes comprising a first mode using a zero biased exponent to represent a biased exponent; a floating-point unit of the programmable processor enabled to perform floating-point operations corresponding to the floating-point instructions and in accordance with floating-point operands and floating-point results each comprising a respective biased exponent, the floating-point operations interpreting the biased exponents of the floating-point operands and producing the biased exponents of the floating-point results in accordance with the current exponent bias and the current zero biased exponent mode; and wherein the instructions further comprise first and second load control resource instructions for respective operations to update the respective state maintained in the exponent bias and zero biased exponent mode control resources.
16 . The apparatus of claim 15 , wherein:
the instructions further comprise a third load control resource instruction, the apparatus further comprises a flush-to-zero control resource of the programmable processor enabled to maintain state representing a current flush-to-zero mode of a plurality of flush-to-zero modes, at least one of the flush-to-zero modes comprising flushing subnormal results to zero, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current flush-to-zero mode.
17 . The apparatus of claim 15 , wherein the zero biased exponent modes further comprise a second mode using the zero biased exponent to represent at least subnormal floating-point numbers.
18 . The apparatus of claim 15 , wherein the programmable processor is one of a plurality of like programmable processors fabricated on a wafer in accordance with wafer-scale integration, and a datacenter element enabled to perform neural network processing comprises the wafer.
19 . The apparatus of claim 15 , wherein the programmable processor comprises one or more hardware registers comprising at least one of the exponent bias and the zero biased exponent mode control resources, and at least one of the first and the second load control resource instructions is a load hardware register instruction.
20 . The apparatus of claim 15 , wherein at least one of the exponent bias and the zero biased exponent mode control resources is memory-mapped, and at least one of the first and the second load control resource instructions is a memory store instruction.
21 . The apparatus of claim 15 , wherein:
the instructions further comprise a third load control resource instruction, the apparatus further comprises a rounding mode control resource of the programmable processor enabled to maintain state representing a current rounding mode of a plurality of rounding modes, at least one of the rounding modes comprising saturating the floating-point results to a value comprising a maximum biased exponent, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current rounding mode.
22 . A method comprising:
maintaining respective state in exponent bias and exponent size control resources of a programmable processor, the respective state representing a current exponent bias of a plurality of exponent biases and a current exponent size of a plurality of exponent sizes; executing instructions by the programmable processor, the instructions comprising floating-point instructions and first and second load control resource instructions; performing floating-point operations by a floating-point unit of the programmable processor, the floating-point operations corresponding to the floating-point instructions and in accordance with floating-point operands and floating-point results each comprising a respective biased exponent, the floating-point operations interpreting the biased exponents of the floating-point operands and producing the biased exponents of the floating-point results in accordance with the current exponent bias and the current exponent size; and updating the state maintained in the exponent bias and exponent size control resources via operations by the first and second load control resource instructions.
23 . The method of claim 22 , wherein:
the exponent size of the exponent size control resource specifies one of a plurality of floating-point formats, a first of the plurality of floating-point formats comprising a biased exponent of N bits and a mantissa of M bits, and a second of the plurality of floating-point formats comprising a biased exponent of (N+delta) bits and a mantissa of (M−delta) bits.
24 . The method of claim 22 , wherein the programmable processor is one of a plurality of like programmable processors fabricated on a wafer in accordance with wafer-scale integration, and a datacenter element enabled to perform neural network processing comprises the wafer.
25 . The method of claim 22 , wherein the programmable processor comprises one or more hardware registers comprising at least one of the exponent bias and the exponent size control resources, and at least one of the first and the second load control resource instructions is a load hardware register instruction.
26 . The method of claim 22 , wherein at least one of the exponent bias and the exponent size control resources is memory-mapped, and at least one of the first and the second load control resource instructions is a memory store instruction.
27 . The method of claim 22 , further comprising:
maintaining rounding mode state in a rounding mode control resource of the programmable processor, the rounding mode state representing a current rounding mode of a plurality of rounding modes, at least one of the rounding modes comprising saturating the floating-point results to a value comprising a maximum biased exponent; and wherein the instructions further comprise a third load control resource instruction, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current rounding mode.
28 . The method of claim 22 , further comprising:
maintaining flush-to-zero mode state in a flush-to-zero mode control resource of the programmable processor, the flush-to-zero mode state representing a current flush-to-zero mode of a plurality of flush-to-zero modes, at least one of the flush-to-zero modes comprising flushing subnormal results to zero; and wherein the instructions further comprise a third load control resource instruction, and the floating-point unit is further enabled to perform the floating-point operations in accordance with the current flush-to-zero mode.
29 . The method of claim 22 , wherein a distinct change in the state maintained in the exponent bias and exponent size control resources is interposed between the sequential executing of each of two floating-point instructions of the floating-point instructions, and wherein the two floating-point instructions are distinct instances of a same floating-point instruction.
30 . The method of claim 22 , wherein a distinct change in the state maintained in the exponent bias and exponent size control resources is interposed between the sequential executing of each of two floating-point instructions of the floating-point instructions, and wherein the two floating-point instructions are respectively responsive to respective fetching of a same instruction from a same memory location.
31 . The method of claim 22 , wherein a distinct change in the state maintained in the exponent bias and exponent size control resources is interposed between the sequential executing of each of two floating-point instructions of the floating-point instructions, and wherein the two floating-point instructions are respectively responsive to respective fetching of identically coded instructions from respective different memory locations.
32 . The method of claim 22 , wherein a distinct change in the state maintained in the exponent bias and exponent size control resources is interposed between the sequential executing of each of two portions of the floating-point instructions, and wherein at least one instruction from each of the two portions has a same instruction opcode value.
33 . The method of claim 22 , wherein a distinct change in the state maintained in the exponent bias and exponent size control resources is interposed between the sequential executing of each of two portions of the floating-point instructions, and wherein at least one instruction from each of the two portions are respective instances of a floating-point multiply-accumulate instruction.
34 . The method of claim 22 , wherein the executing of one of the first and second load control resource instructions is one of: unconditional, non-selective, static, and a-priori determined.
35 . The method of claim 22 , wherein the executing of one of the first and second load control resource instructions is one of: conditional, selective, dynamic, and not a-priori determined.Join the waitlist — get patent alerts
Track US2022172030A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.