US2023064381A1PendingUtilityA1

Memory-Size- and Bandwidth-Efficient Method for Feeding Systolic Array Matrix Multipliers

Assignee: INTEL CORPPriority: Jul 7, 2017Filed: May 9, 2022Published: Mar 2, 2023
Est. expiryJul 7, 2037(~10.9 yrs left)· nominal 20-yr term from priority
G06F 2207/3892G06F 7/5443G06F 17/16
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Matrix multiplication systolic array feed methods and related processing element (PE) microarchitectures for efficiently implementing systolic array generic matrix multiplier (SGEMM) in integrated circuits is provided. A systolic array architecture may include a processing element array, a column feeder array, and a row feeder array. A bandwidth of external memory may be reduced by a factor of reduction based on interleaving of the matrix data via a feeding pattern of the column feeder array and the row feeder array.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . An integrated circuit comprising:
 accumulator control circuitry configurable to:
 modify first matrix data of a first matrix based on a control signal, wherein the first matrix is a sparse matrix having one or more zeros; and 
 send modified first matrix data and second matrix data of a second matrix to feeder circuitry; 
   the feeder circuitry configurable to feed the modified first matrix data and the second matrix data into one or more processing elements to facilitate zero skipping; and   the one or more processing elements configurable to:
 receive the modified first matrix data and the second matrix data; and 
 perform sparse matrix multiply-accumulate operations using the modified first matrix data and the second matrix data, wherein the sparse matrix multiply-accumulate operations comprise zero skipping based at least in part on the control signal. 
   
     
     
         3 . The integrated circuit of  claim 2 , wherein the one or more processing elements are configurable to perform fused multiply-add (FMA) operations when performing the sparse matrix multiply-accumulate operations. 
     
     
         4 . The integrated circuit of  claim 3 , wherein the control signal indicates which fused multiply-add operations to perform and which fused multiply-add operations to skip to facilitate zero skipping. 
     
     
         5 . The integrated circuit of  claim 2 , wherein the one or more processing elements are configurable to perform fewer fused multiply-add operations when performing the sparse matrix multiply-accumulate operations than when performing matrix multiply-accumulate operations when the first matrix does not comprise the sparse matrix. 
     
     
         6 . The integrated circuit of  claim 2 , wherein the control signal comprises an index associated with the sparse matrix. 
     
     
         7 . The integrated circuit of  claim 2 , wherein the accumulator control circuitry facilitates zero-skipping with a schedule of non-zero values to be fed into the one or more processing elements. 
     
     
         8 . An integrated circuit to perform sparse matrix multiplication that avoids multiplying zeros of a sparse matrix, comprising:
 feeder circuitry to feed first matrix data of a first matrix and second matrix data of a second matrix into one or more processing elements, wherein the first matrix is a sparse matrix comprising a plurality of zeros; and   the one or more processing elements to receive the first matrix data, the second matrix data, and the control signal and perform sparse matrix multiply-accumulate operations using the first matrix data and the second matrix data that skip multiplying zeros of the first matrix.   
     
     
         9 . The integrated circuit of  claim 8 , wherein the feeder circuitry provides a control signal that signals to the one or more processing elements to perform the sparse matrix multiply-accumulate operations that skip multiplying the zeros of the first matrix. 
     
     
         10 . The integrated circuit of  claim 9 , wherein the one or more processing elements are to perform fused multiply-add (FMA) operations when performing the sparse matrix multiply-accumulate operations. 
     
     
         11 . The integrated circuit of  claim 10 , wherein the control signal indicates which fused multiply-add operations to perform and which fused multiply-add operations to skip to facilitate zero skipping. 
     
     
         12 . The integrated circuit of  claim 10 , wherein the one or more processing elements are configurable to perform fewer fused multiply-add operations when performing the sparse matrix multiply-accumulate operations than when performing matrix multiply-accumulate operations when the first matrix does not comprise the sparse matrix. 
     
     
         13 . The integrated circuit of  claim 10 , wherein the control signal comprises an index associated with the sparse matrix. 
     
     
         14 . The integrated circuit of  claim 13 , wherein the control signal indicates a schedule of non-zero values to be fed into the one or more processing elements. 
     
     
         15 . A method comprising:
 providing, to one or more processing elements, first sparse matrix data corresponding to a first sparse matrix that comprises zeros;   providing, to the one or more processing elements, second matrix data corresponding to a second matrix;   providing, to the one or more processing elements, an index control signal that corresponds to the zeros of the first matrix; and   using the one or more processing elements to perform sparse matrix multiplication based on the first sparse matrix data, the second matrix data, and the index control signal.   
     
     
         16 . The method of  claim 15 , wherein performing sparse matrix multiplication comprises performing one or more fused multiply-add (FMA) operations. 
     
     
         17 . The method of  claim 16 , wherein the index control signal indicates which fused multiply-add operations to perform and which fused multiply-add operations to skip to facilitate zero skipping. 
     
     
         18 . The method of  claim 16 , wherein the index control signal indicates a schedule of non-zero values of the first sparse matrix data to be fed into the one or more processing elements. 
     
     
         19 . The method of  claim 15 , wherein the first sparse matrix data is provided by row feeder circuitry and the second matrix data is provided by column feeder circuitry. 
     
     
         20 . The method of  claim 15 , wherein the method is performed in the recited order. 
     
     
         21 . The method of  claim 15 , comprising modifying the first sparse matrix or the second matrix to facilitate zero skipping in the one or more processing elements.

Join the waitlist — get patent alerts

Track US2023064381A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.