US2023359695A1PendingUtilityA1

Memory-Size- and Bandwidth-Efficient Method for Feeding Systolic Array Matrix Multipliers

Assignee: INTEL CORPPriority: Jul 7, 2017Filed: Jul 17, 2023Published: Nov 9, 2023
Est. expiryJul 7, 2037(~10.9 yrs left)· nominal 20-yr term from priority
G06F 17/16G06F 7/5443G06F 2207/3892
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Matrix multiplication systolic array feed methods and related processing element (PE) microarchitectures for efficiently implementing systolic array generic matrix multiplier (SGEMM) in integrated circuits is provided. A systolic array architecture may include a processing element array, a column feeder array, and a row feeder array. A bandwidth of external memory may be reduced by a factor of reduction based on interleaving of the matrix data via a feeding pattern of the column feeder array and the row feeder array.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . Circuitry of an integrated circuit, comprising:
 a plurality of processing elements comprising respective dot product circuitry and accumulator circuitry;   loading circuitry to retrieve a first block of a first matrix and a first block of a second matrix from a memory; and   feeder circuitry comprising a plurality of banks to store respective segments of the first block of the first matrix and respective segments of the first block of the second matrix and to feed the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time.   
     
     
         3 . The circuitry of  claim 2 , wherein the respective accumulator circuitry of the plurality of processing elements is controllable to be drained upon accumulating results of dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix obtained using the respective dot product circuitry of the plurality of processing elements. 
     
     
         4 . The circuitry of  claim 3 , wherein the accumulated results of the dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix correspond to a dot product of the first block of the first matrix and the first block of the second matrix. 
     
     
         5 . The circuitry of  claim 2 , wherein the plurality of processing elements are arranged in a systolic array. 
     
     
         6 . The circuitry of  claim 2 , wherein the feeder circuitry comprises:
 an array of column feeders respectively:
 coupled to a respective processing element of an outermost row of the plurality of processing elements; and 
 comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another column feeder; and 
   an array of row feeders respectively:
 coupled to a respective processing element of an outermost column of the plurality of processing elements; and 
 comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another row feeder. 
   
     
     
         7 . The circuitry of  claim 6 , wherein feeding the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time comprises:
 the array of column feeders feeding a first segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix; and   the array of column feeders feeding a second segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix.   
     
     
         8 . The circuitry of  claim 2 , wherein the first block of the first matrix corresponds to a column of the first matrix and the first block of the second matrix corresponds to a row of the second matrix. 
     
     
         9 . A method comprising:
 loading a first block of a first matrix and a first block of a second matrix from memory into feeder circuitry;   feeding different segments of the first block of the first matrix and the first block of the second matrix at different times from the feeder circuitry into a plurality of processing elements; and   obtaining a matrix multiplication product of the first block of the first matrix and the first block of the second matrix by accumulating matrix multiplications of the different segments of the first block of the first matrix and the first block of the second matrix at the different times using the plurality of processing elements.   
     
     
         10 . The method of  claim 9 , wherein feeding the different segments of the first block of the first matrix and the first block of the second matrix and obtaining the matrix multiplication product comprises:
 feeding, into the plurality of processing elements, a first segment of a plurality of segments of the first block of the first matrix to the plurality of processing elements and feeding a first segment of a plurality of segments of the first block of the second matrix to the plurality of processing elements;   performing a first matrix multiplication of the first segment of the first block of the first matrix and the first segment of the first block of the second matrix using the plurality of processing elements at a first time;   feeding, into the plurality of processing elements, a second segment of a plurality of segments of the first block of the first matrix to the plurality of processing elements and feeding the first segment of a plurality of segments of the first block of the second matrix to the plurality of processing elements;   performing a second matrix multiplication of the second segment of the first block of the first matrix and the first segment of the first block of the second matrix using the plurality of processing elements at a second time; and   accumulating results of the first matrix multiplication and the second matrix multiplication.   
     
     
         11 . The method of  claim 9 , wherein the first block comprises a first row of the first matrix and the second block comprises a first column of the second matrix. 
     
     
         12 . The method of  claim 9 , wherein obtaining the matrix multiplication product comprises performing a plurality of dot product operations in the plurality of processing elements arranged in a systolic array. 
     
     
         13 . The method of  claim 12 , wherein obtaining the matrix multiplication product comprises accumulating results of the plurality of dot product operations in the plurality of processing elements. 
     
     
         14 . An article of manufacture comprising tangible, non-transitory, machine-readable media comprising instructions to implement the following circuitry on a programmable logic device:
 a plurality of processing elements comprising respective dot product circuitry and accumulator circuitry;   loading circuitry to retrieve a first block of a first matrix and a first block of a second matrix from a memory; and   feeder circuitry comprising a plurality of banks to store respective segments of the first block of the first matrix and respective segments of the first block of the second matrix and to feed the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time.   
     
     
         15 . The article of manufacture of  claim 14 , wherein the respective accumulator circuitry of the plurality of processing elements is controllable to be drained upon accumulating results of dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix obtained using the respective dot product circuitry of the plurality of processing elements. 
     
     
         16 . The article of manufacture of  claim 15 , wherein the accumulated results of the dot products of the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix correspond to a dot product of the first block of the first matrix and the first block of the second matrix. 
     
     
         17 . The article of manufacture of  claim 14 , wherein the plurality of processing elements are arranged in a systolic array. 
     
     
         18 . The article of manufacture of  claim 14 , wherein the feeder circuitry comprises:
 an array of column feeders respectively:
 coupled to a respective processing element of an outermost row of the plurality of processing elements; and 
 comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another column feeder; and 
   an array of row feeders respectively:
 coupled to a respective processing element of an outermost column of the plurality of processing elements; and 
 comprising a respective buffer memory to serve as a bank to temporarily store data before it is provided to the respective processing element or to another row feeder. 
   
     
     
         19 . The article of manufacture of  claim 18 , wherein feeding the respective segments of the first block of the first matrix and the respective segments of the first block of the second matrix to the plurality of processing elements interleaved over time comprises:
 the array of column feeders feeding a first segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix; and   the array of column feeders feeding a second segment of the respective segments of the first block of the first matrix as the array of row feeders feeds the respective segments of the first block of the second matrix.   
     
     
         20 . The article of manufacture of  claim 14 , wherein the first block of the first matrix corresponds to a column of the first matrix and the first block of the second matrix corresponds to a row of the second matrix. 
     
     
         21 . The article of manufacture of  claim 14 , wherein the circuitry is configured to be operated according to a method comprising:
 loading the first block of the first matrix and the first block of the second matrix from the memory into the feeder circuitry using the loading circuitry;

Join the waitlist — get patent alerts

Track US2023359695A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.