Using a low-bit-width dot product engine to sum high-bit-width numbers
Abstract
A system includes a vector multiplier configured to multiply a first vector of integer elements with a second vector of integer elements to determine a resulting vector of integer elements, wherein integer elements of the first and second vectors of integer elements are represented using a first number of bits and an integer element of the first vector of integer elements represents a portion of a value of a group of values. The system further includes a vector adder configured to add together the integer elements of the resulting vector of integer elements to determine a summed result, a bit shifter configured to shift bits of the summed result leftward, and an accumulator configured to determine an accumulated output sum that includes the leftward-shifted summed result.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system, comprising:
a vector multiplier configured to multiply a first vector of integer elements with a second vector of integer elements to determine a resulting vector of integer elements, wherein:
integer elements of the first and second vectors of integer elements are represented using a first number of bits; and
an integer element of the first vector of integer elements represents a corresponding portion of a corresponding value of a group of values, wherein:
a value of the group of values is represented using a second number of bits greater than the first number of bits; and
the value of the group of values is stored as split segments across more than one integer element of the first vector of integer elements;
a vector adder configured to add together the integer elements of the resulting vector of integer elements to determine a summed result; a bit shifter configured to shift bits of the summed result leftward; and an accumulator configured to determine an accumulated output sum that includes the leftward-shifted summed result.
2 . The system of claim 1 , wherein each integer element of the second vector of integer elements has a value that is either zero or one.
3 . The system of claim 1 , wherein the first number of bits is eight bits and the second number of bits is thirty-two bits.
4 . The system of claim 1 , further comprising a multiplexer configured to select a bit shift amount of the bit shifter based at least in part on an iteration count.
5 . The system of claim 1 , further comprising a first storage unit of one or more registers configured to store the first vector of integer elements.
6 . The system of claim 5 , wherein the first storage unit is configured to store at least 256 bits.
7 . The system of claim 1 , further comprising a second storage unit of one or more registers configured to store the second vector of integer elements.
8 . The system of claim 1 , wherein the accumulated output sum includes a plurality of leftward-shifted summed results.
9 . The system of claim 1 , wherein the accumulated output sum is utilized in an artificial neural network operation.
10 . A system, comprising:
a vector multiplier configured to multiply a first vector of floating-point elements with a second vector of floating-point elements to determine a resulting vector of floating-point elements, wherein:
floating-point elements of the first and second vectors of floating-point elements are represented using a first number of bits; and
a floating-point element of the first vector of floating-point elements represents a corresponding portion of a corresponding value of a group of values, wherein:
a value of the group of values is represented using a second number of bits greater than the first number of bits; and
the value of the group of values is stored as split segments across more than one floating-point element of the first vector of floating-point elements;
a vector adder configured to add together the floating-point elements of the resulting vector of floating-point elements to determine a summed result; a subtractor configured to subtract a subtraction amount from an exponent portion of the summed result to determine an exponent-modified result; and an accumulator configured to determine an accumulated output sum that includes the exponent-modified result.
11 . The system of claim 10 , wherein each floating-point element of the second vector of floating-point elements has a value that is either zero or one.
12 . The system of claim 10 , wherein the first number of bits is sixteen bits and the second number of bits is thirty-two bits.
13 . The system of claim 12 , wherein the first number of bits are formatted in a Brain Floating Point floating-point format and the second number of bits are formatted in a single-precision floating-point format.
14 . The system of claim 10 , further comprising a multiplexer configured to select the subtraction amount of the subtractor based at least in part on an iteration count.
15 . The system of claim 10 , wherein the subtractor comprises an adder configured to add negative numbers.
16 . The system of claim 10 , further comprising a first storage unit of one or more registers configured to store the first vector of floating-point elements.
17 . The system of claim 16 , wherein the first storage unit is configured to store at least 256 bits.
18 . The system of claim 10 , further comprising a second storage unit of one or more registers configured to store the second vector of floating-point elements.
19 . The system of claim 10 , wherein the accumulated output sum includes a plurality of exponent-modified results.
20 . The system of claim 10 , wherein the accumulated output sum is utilized in an artificial neural network operation.Join the waitlist — get patent alerts
Track US2023056304A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.