US2023359695A1PendingUtilityA1
Memory-Size- and Bandwidth-Efficient Method for Feeding Systolic Array Matrix Multipliers
Est. expiryJul 7, 2037(~10.9 yrs left)· nominal 20-yr term from priority
Inventors:Jack Z. YingerAndrew Chaang LingTomasz CzajkowskiDavor CapalijaEriko NurvitadhiDeborah T. Marr
G06F 17/16G06F 7/5443G06F 2207/3892
71
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Matrix multiplication systolic array feed methods and related processing element (PE) microarchitectures for efficiently implementing systolic array generic matrix multiplier (SGEMM) in integrated circuits is provided. A systolic array architecture may include a processing element array, a column feeder array, and a row feeder array. A bandwidth of external memory may be reduced by a factor of reduction based on interleaving of the matrix data via a feeding pattern of the column feeder array and the row feeder array.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . Circuitry of an integrated circuit, comprising:
a plurality of processing elements comprising respective dot product circuitry and accumulator circuitry; loading circuitry to retrieve a first block of a first matrix and a first block of a second matrix from a memory; and feeder circuitry comprising a plurality of banks to store respective segments of the first block of the first matrix and respective segments of the first block of the second matrix and to feed the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time.
3 . The circuitry of claim 2 , wherein the respective accumulator circuitry of the plurality of processing elements is controllable to be drained upon accumulating results of dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix obtained using the respective dot product circuitry of the plurality of processing elements.
4 . The circuitry of claim 3 , wherein the accumulated results of the dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix correspond to a dot product of the first block of the first matrix and the first block of the second matrix.
5 . The circuitry of claim 2 , wherein the plurality of processing elements are arranged in a systolic array.
6 . The circuitry of claim 2 , wherein the feeder circuitry comprises:
an array of column feeders respectively:
coupled to a respective processing element of an outermost row of the plurality of processing elements; and
comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another column feeder; and
an array of row feeders respectively:
coupled to a respective processing element of an outermost column of the plurality of processing elements; and
comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another row feeder.
7 . The circuitry of claim 6 , wherein feeding the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time comprises:
the array of column feeders feeding a first segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix; and the array of column feeders feeding a second segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix.
8 . The circuitry of claim 2 , wherein the first block of the first matrix corresponds to a column of the first matrix and the first block of the second matrix corresponds to a row of the second matrix.
9 . A method comprising:
loading a first block of a first matrix and a first block of a second matrix from memory into feeder circuitry; feeding different segments of the first block of the first matrix and the first block of the second matrix at different times from the feeder circuitry into a plurality of processing elements; and obtaining a matrix multiplication product of the first block of the first matrix and the first block of the second matrix by accumulating matrix multiplications of the different segments of the first block of the first matrix and the first block of the second matrix at the different times using the plurality of processing elements.
10 . The method of claim 9 , wherein feeding the different segments of the first block of the first matrix and the first block of the second matrix and obtaining the matrix multiplication product comprises:
feeding, into the plurality of processing elements, a first segment of a plurality of segments of the first block of the first matrix to the plurality of processing elements and feeding a first segment of a plurality of segments of the first block of the second matrix to the plurality of processing elements; performing a first matrix multiplication of the first segment of the first block of the first matrix and the first segment of the first block of the second matrix using the plurality of processing elements at a first time; feeding, into the plurality of processing elements, a second segment of a plurality of segments of the first block of the first matrix to the plurality of processing elements and feeding the first segment of a plurality of segments of the first block of the second matrix to the plurality of processing elements; performing a second matrix multiplication of the second segment of the first block of the first matrix and the first segment of the first block of the second matrix using the plurality of processing elements at a second time; and accumulating results of the first matrix multiplication and the second matrix multiplication.
11 . The method of claim 9 , wherein the first block comprises a first row of the first matrix and the second block comprises a first column of the second matrix.
12 . The method of claim 9 , wherein obtaining the matrix multiplication product comprises performing a plurality of dot product operations in the plurality of processing elements arranged in a systolic array.
13 . The method of claim 12 , wherein obtaining the matrix multiplication product comprises accumulating results of the plurality of dot product operations in the plurality of processing elements.
14 . An article of manufacture comprising tangible, non-transitory, machine-readable media comprising instructions to implement the following circuitry on a programmable logic device:
a plurality of processing elements comprising respective dot product circuitry and accumulator circuitry; loading circuitry to retrieve a first block of a first matrix and a first block of a second matrix from a memory; and feeder circuitry comprising a plurality of banks to store respective segments of the first block of the first matrix and respective segments of the first block of the second matrix and to feed the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time.
15 . The article of manufacture of claim 14 , wherein the respective accumulator circuitry of the plurality of processing elements is controllable to be drained upon accumulating results of dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix obtained using the respective dot product circuitry of the plurality of processing elements.
16 . The article of manufacture of claim 15 , wherein the accumulated results of the dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix correspond to a dot product of the first block of the first matrix and the first block of the second matrix.
17 . The article of manufacture of claim 14 , wherein the plurality of processing elements are arranged in a systolic array.
18 . The article of manufacture of claim 14 , wherein the feeder circuitry comprises:
an array of column feeders respectively:
coupled to a respective processing element of an outermost row of the plurality of processing elements; and
comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another column feeder; and
an array of row feeders respectively:
coupled to a respective processing element of an outermost column of the plurality of processing elements; and
comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another row feeder.
19 . The article of manufacture of claim 18 , wherein feeding the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time comprises:
the array of column feeders feeding a first segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix; and the array of column feeders feeding a second segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix.
20 . The article of manufacture of claim 14 , wherein the first block of the first matrix corresponds to a column of the first matrix and the first block of the second matrix corresponds to a row of the second matrix.
21 . The article of manufacture of claim 14 , wherein the circuitry is configured to be operated according to a method comprising:
loading the first block of the first matrix and the first block of the second matrix from the memory into the feeder circuitry using the loading circuitry;Join the waitlist — get patent alerts
Track US2023359695A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.