Power saving floating point multiplier-accumulator with precision-aware accumulation
Abstract
A floating point multiplier-accumulator (MAC) multiplies and accumulates N pairs of floating point values using N MAC processors operating simultaneously, each pair of values comprising an input value and a coefficient value to be multiplied and accumulated. The pairs of floating point values are simultaneously processed by the plurality of MAC processors, each of which outputs a signed integer form fraction and a maximum exponent. A range estimator forms a possible range of values from the exponent differences and determines an adder precision. The integer form fractions are summed using the adder precision, a sign bit is extracted, and a floating point value is output. Each MAC processor provides its integer form fraction with a precision determined by the MAC processor's exponent difference.
Claims
exact text as granted — not AI-modifiedI claim:
1. A floating point multiplier-accumulator (MAC) multiplying and accumulating N pairs of values, each pair of values comprising a respective input value and a corresponding coefficient value, the floating point MAC comprising:
a plurality N of MAC processors, each MAC processor receiving the respective input value and the corresponding coefficient value, each MAC processor comprising:
a sign processor configured to perform an exclusive OR operation on a sign bit of the respective input value and a sign bit of the corresponding coefficient value, the sign processor outputting a corresponding sign bit;
a mantissa processor configured to perform an integer multiplication of a mantissa of the respective input value and a mantissa of the corresponding coefficient value and outputting a fraction;
an exponent processor determining an exponent sum of an exponent of the respective input value and an exponent of the corresponding coefficient value, the exponent processor receiving a maximum exponent from a centralized find maximum exponent processor, the exponent processor modifying the maximum exponent and also outputting an exponent difference between the maximum exponent and the exponent sum;
a Pad, Complement, Shift (PCS) Processor receiving the fraction from the mantissa processor, the corresponding sign bit from the sign processor, and the exponent difference from the exponent processor, the PCS processor configured to pad the fraction by pre-pending and appending 0s to the fraction to generate a first value, thereafter performing a two's complement of the first value if the corresponding sign bit from the sign processor is negative and otherwise taking no action on the first value to generate a second value, the PCS processor configured to performing a shift operation on the second value by right shifting the second value by the exponent difference to generate a PCS output;
the centralized find maximum exponent processor receiving an exponent sum from each exponent processor of the MAC processors, the centralized find maximum exponent processor outputting a maximum exponent value corresponding to a maximum exponent sum;
a binary tree of adders summing N PCS output values to a single value;
a final stage normalizing the single value, generating a final stage mantissa by performing a 2s complement if the single value is negative, generating a final stage sign bit, and concatenating the final stage sign bit, final stage mantissa, and maximum exponent into a floating point MAC result.
2. The floating point MAC of claim 1 where the exponent processor computes a mantissa precision for the mantissa processor of a MAC processor based on the exponent difference.
3. The floating point MAC of claim 2 where the mantissa precision is 4 bits when the exponent difference is greater than 24.
4. The floating point MAC of claim 2 where the mantissa precision is 8 bits when the exponent difference is greater than 21.
5. The floating point MAC of claim 1 where the mantissa precision is 12 bits when the exponent difference is larger than 12.
6. The floating point MAC of claim 1 where the exponent difference of a MAC processor that does not have the maximum exponent sum is incremented if the mantissa processor does not overflow and the exponent difference is 0.
7. The floating point MAC of claim 1 where the exponent difference of a MAC processor that does not have the maximum exponent sum is decremented if the mantissa processor has an overflow and the exponent difference is greater than 0.
8. The floating point MAC of claim 1 where the maximum exponent is incremented if the exponent difference of a MAC processor is 0 and an associated mantissa processor has a multiplication overflow.
9. The floating point MAC of claim 1 where each MAC processor exponent processor performs an estimate of minimum value and maximum value based on an associated exponent difference.
10. The floating point MAC of claim 1 where the binary tree of adders has a variable precision.
11. The floating point MAC of claim 10 where a sum of exponent processor minimum values and a sum of exponent processor maximum values determines a particular precision of the variable precision of the binary tree of adders.
12. The floating point MAC of claim 11 where the binary tree of adders has a full precision and less than full precision, and the less than full precision is enabled when the sum of exponent processor maximum values and the sum of exponent processor minimum values are either both positive values or both negative values.
13. The floating point MAC of claim 11 where each adder of the binary tree of adders comprises cascaded 8 bit adders.
14. The floating point MAC of claim 11 where at least one adder of the binary tree of adders is selectively configurable in a 16 bit mode, a 24 bit mode, and a 32 bit mode.
15. A floating point multiplier-accumulator (MAC) multiplying and accumulating N pairs comprising an input value and a coefficient value, the floating point MAC comprising:
a plurality N of MAC processors, each MAC processor receiving a respective input value and a corresponding coefficient value, each MAC processor comprising:
a sign processor configured to perform an exclusive OR operation on a sign bit of the respective input value and a sign bit of the corresponding coefficient value and outputting a sign bit;
a mantissa processor configured to perform an integer multiplication of a hidden bit restored mantissa of the respective input value and a hidden bit restored mantissa of the corresponding coefficient value and output a fraction, the mantissa processor dividing the output fraction by two and asserting an exponent increment upon an overflow condition;
an exponent processor generating an exponent sum of an exponent of the respective input value and an exponent of the corresponding coefficient value, the exponent processor receiving a maximum exponent from a centralized find maximum exponent processor, the exponent processor modifying the maximum exponent and also outputting an exponent difference computed by subtracting the exponent sum from the maximum exponent, the exponent processor also using the exponent difference and sign bit to estimate a minimum value and a maximum value;
a Pad, Complement, Shift (PCS) Processor receiving the fraction from the mantissa processor and also the sign bit from the sign processor, the PCS processor configured to pad the fraction by pre-pending and appending 0s to the fraction to generate a first value, thereafter generating a second value by performing a two's complement of the first value if the sign bit is negative and otherwise taking no action on the first value, the PCS processor configured to performing a shift operation on the second value by right shifting the second value by the exponent difference to generate a PCS output;
the centralized find maximum exponent processor receiving an exponent sum from each MAC processor exponent processor, the centralized find maximum exponent processor outputting a maximum exponent value corresponding to a maximum exponent processor sum;
a central range estimator configured to sum minimum values from the MAC processor exponent processors and also to sum maximum values from the MAC processor exponent processors, the central range estimator outputting an adder precision based on the sum of minimum values and the sum of maximum values;
a binary tree of adders summing N PCS output values to a single value, the adders configured to sum using the adder precision of the central range estimator;
a final stage normalizing the single value, generating a final stage sign bit from the single value, generating a final stage mantissa by performing a 2s complement of the single value if the final stage sign bit is negative, and concatenating the final stage sign bit, final stage mantissa, and an adjusted maximum exponent into a MAC result.
16. The floating point MAC of claim 15 where each MAC processor exponent processor computes a mantissa precision based on the exponent difference.
17. The floating point MAC of claim 16 where the mantissa precision is 4 bits when the exponent difference is greater than 24.
18. The floating point MAC of claim 16 where the mantissa precision is 8 bits when the exponent difference is greater than 21.
19. The floating point MAC of claim 16 where the mantissa precision is 12 bits when the exponent difference is larger than 12.
20. The floating point MAC of claim 16 where the adjusted maximum exponent is 8 bits and is equal to the maximum exponent less a number of leading 0s of the single value which exceed a number of prepended 0s less 127.Join the waitlist — get patent alerts
Track US12106069B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.